Evaluating a commission: measuring the value of public art honestly

Public art benefits split into two lists usually printed as one: five things that can be counted with reasonable honesty, and a much longer set that is asserted, believed and never tested. Evaluating a commission means keeping them apart. The measurable list is short and unglamorous: how many people pass the work, how long they stop, what a sample say, what the press and online record shows, and what it costs to maintain against forecast. The unmeasurable list holds the two claims funding applications make most often: that the work is good art, and that it regenerated the area.
Most evaluations are worthless not through dishonesty but because nobody collected a baseline. A footfall count taken after installation, with nothing to compare against, is a number rather than a finding.
The only fair test: did the brief’s stated purpose happen
The brief’s stated purpose is the only test a commission can fairly be held to: the one the artist agreed to answer and the money was approved against. A brief saying the work should mark the entrance to a redeveloped quarter and be legible from the far end of the street sets a test checkable by standing at the far end of the street. A brief saying the work should be iconic, celebrate heritage and create a sense of place sets no test at all, and no evaluation of it will be more than an opinion piece with photographs.
Evaluation therefore belongs in the brief rather than the closing report. Writing the purpose as something observable, then agreeing what evidence would show it had happened, converts an argument at the end into a measurement. It also protects the artist, who is otherwise judged against criteria invented once the work was in the ground.
What can be counted, and what each count cannot tell you
Counting is possible in five areas, and each carries a limit that belongs in the same breath as the number, because a measure quoted without its caveat starts the problem.
| Measure | Method | What it cannot tell you |
|---|---|---|
| Footfall | Counters or manual counts at fixed points and times, before installation and after, with a control location counted the same days | Whether any change came from the artwork rather than the new shopfront, the resurfaced footway or the season. Without a control site it tells you almost nothing. |
| Dwell time | Timed observation of a sample of passers by: how many stop, for how long, and what they do | Why they stopped, and whether stopping is a benefit. A work people stop at because it blocks the desire line scores well and is a failure. |
| Survey response | Intercept surveys on site plus an online panel, with a fixed question set repeated so results stay comparable | The view of everyone who avoided the site. Around 400 responses gives roughly plus or minus 5 percentage points at 95 per cent confidence for a random sample; an intercept survey is not random, so that is a floor on the error. |
| Media and online record | Counts of press, broadcast and social items, coded positive, negative or neutral, local or national | What the coverage did. Advertising value equivalence, converting column inches into cash, is rejected as a measure of value by the communications measurement field. |
| Maintenance cost against forecast | Actual annual spend on inspection, cleaning, repair and repainting, against the original budget figure | Nothing about the artistic outcome, everything about the quality of the commissioning. It is the most useful and least used measure available. |
What cannot be measured, and saying so
Two questions are asked of every finished commission and neither can be answered by evaluation. The first is whether the work is any good. Aesthetic quality is contested by design, judgements change across a generation, and the works provoking the strongest early hostility are not reliably the bad ones. Counting positive survey responses measures acceptance, and calling that quality is a category error.
The second is whether the work regenerated anything. A commission arrives alongside new paving, new lighting, a transport change, several lease decisions and a national economic cycle, and no method available to a project budget separates its contribution from the rest. The counterfactual, what would have happened without the artwork, is not observable. Where a work is the only intervention in an unchanged street, a before and after comparison with a control street is worth doing, and remains weak evidence rather than proof.
Economic impact claims, and an abstention this page keeps
Economic impact claims are the currency of public art advocacy and are almost always built badly. The public sector owns the vocabulary for what is wrong with them, in the appraisal guidance government uses on itself: additionality, the share of the outcome that would not have happened anyway; deadweight, the share that would; displacement, activity moved rather than created; and leakage, benefit flowing outside the area counted. A claim that a work generated so many visitors and so much spend, with none of those four applied, is a gross count presented as a net one. Two further failures are routine: ratios of the form “every pound spent returned several pounds” are usually expenditure multipliers applied to construction spend, equally true of a car park; and figures lifted from another country come from different funding and tax systems.
So this page offers no return figure for public art, and no other page here will supply one. That is an abstention rather than a hedge: the number does not exist in a form that survives scrutiny, and reprinting somebody else’s is worth less than saying nothing. What can be said is narrower. Spending is traceable, so the share reaching artists, fabricators and contractors, and how much stayed inside a defined travel to work area, is reportable as fact. Maintenance cost is knowable. Whether the commission delivered what the brief asked is knowable. Those three findings serve the next committee better than a multiplier nobody can defend.
Designing an evaluation that will actually happen
Designing an evaluation means fixing five things before the artist is appointed, since each becomes impossible or dishonest afterwards.
- State the purpose in observable terms, and write down what evidence would show it was met.
- Collect the baseline before anything changes on site: footfall counts, fixed viewpoint photographs, a first survey round.
- Name the evaluator and the budget line, since a client evaluating its own commission produces a document nobody outside believes.
- Set reporting points at handover, twelve months and three to five years: the interesting findings are the late ones.
- Agree in advance that a negative finding will be published, since a regime that can only produce good news is marketing.
Proportion matters. A commission of a few thousand pounds does not warrant a formal study, and an honest one page record of what was spent, what was made, what it costs to keep and what people said beats a consultant’s report. A large permanent commission justifies a baseline, a control location and repeat measurement over years. A programme’s most valuable output is rarely about a single work: it is the pattern across a dozen, showing which briefs produced usable proposals, which materials cost what to maintain, and which sites turned out harder than predicted.