Moving beyond metric capture
Goodhardt's law is a fixture among a certain set of wonks, techies, and internet native types.
the moment a measure becomes a target, it ceases to be a good measure.
Cobras offer a canonical example. Cobras you ask. Yes in the British raj coras presented a problem so there was a bounty placed on cobras. As a result cobras were breed. When the authorities discovered that, they ended the bounty program and the breeders let their cobras free, resulting in more cobras than at the start of the episode.
One experiment I've long been curious to run is moving from purely quantative metrics (such as number of cobras killed) to rigorous engagement with the underlying question (such as are cobras presenting a problem for people).
Moody's doesn't publish a one-time assessment of a municipality's fiscal health and call it done. It maintains a living judgment, updated as conditions change, sensitive to leading indicators, explicit about what it's watching. Downgrades happen mid-cycle. Outlooks shift from stable to negative before anything catastrophic occurs. The whole apparatus is designed to surface trouble while there's still time to act, and the market pays for that signal because the market moves on it.
Imagine a parallel apparatus for public programs: not an annual performance report but an ongoing impact rating that synthesizes administrative data, community feedback, and independent analysis into a regularly updated assessment. "Watch status." "Outlook: deteriorating." "Core outcomes stable; equity indicators warrant review." The language of credit markets applied to the public interest, with AI doing the continuous synthesis that makes frequent updates tractable rather than prohibitively expensive.
Consider how this would work for something as ordinary as a new after-school program in a mid-size school district. Under the current model, the district applies for a grant, defines success as "200 students enrolled" and "10% improvement in reading scores," runs the program for three years, and submits a final report. Maybe the program is working brilliantly for third graders and failing completely for sixth graders. Maybe attendance craters after the first semester because the bus schedule changed. Maybe the most important thing happening is that a cluster of kids who were otherwise unsupervised between 3 and 6 PM are no longer showing up in juvenile incident reports. None of this nuance survives the binary pass/fail of the final evaluation. A living impact rating would catch the bus schedule problem in month four. It would flag the divergence between age cohorts. It would notice the juvenile incident correlation before anyone thought to look for it, because a well-designed rating system watches for what's actually happening, not just what the original grant application promised would happen.
Source of the long quote is a blog post of mine from a while back: https://pioneeringspirit.xyz/outcomes-on-tap
Context on the cobra effect: https://www.historic-uk.com/HistoryUK/HistoryofBritain/Cobra-Effect/