The Measurement Trap: Why Optimizing the Score Breaks the Thing You Were Scoring

A marketing team A/B tests subject lines, measures open rates, iterates on the winner, runs the next test. Three months in: open rates are up 22%. Reply rate is down. They optimized the score. The score broke.

A target bullseye with concentric orange rings on a cream background, with a navy arrow striking the outer ring rather than the center, representing Goodhart's Law in copy optimization.

A B2B marketing team runs a disciplined testing program. Six months of A/B tests on subject lines, iterating on the winners, checking open rates, running the next test. They have the spreadsheet. They have the system. They are doing exactly what every growth playbook says to do.

Their open rates are up 22%. Their reply rate is down 30%.

Nobody in the post-mortem can explain it, because the work looks right. Every test was valid. Every winning variant was measured correctly. The problem is not the execution. The problem is that they optimized for a metric, and the metric stopped measuring the thing they actually cared about.

This is the measurement trap. Not “metrics are bad” and not “data-driven marketing is wrong.” Something more specific: once you make a metric the target of your optimization, behavior shifts toward producing the metric, and the metric quietly stops being a reliable measure of the underlying thing it was supposed to track.

It has a name. In economics it’s called Goodhart’s Law, named for Bank of England adviser Charles Goodhart: “When a measure becomes a target, it ceases to be a good measure.” The phenomenon was originally observed in monetary policy, but the structure appears in every complex system where the proxy is easier to optimize than the goal it stands in for: corporate performance management, machine learning reward functions, standardized testing, and B2B marketing copy.


Why this happens structurally

The underlying goal in B2B copy is something like: connect with this specific buyer, at this stage in their consideration process, in a way that advances their confidence in the fit between your product and their problem. That’s what you’re trying to produce.

That goal is difficult to measure directly. You can’t score “confidence in fit” after every email. So you measure proxies: open rate, click rate, reply rate, meeting booked, stage advancement, content engagement score.

Each proxy has a gap to the underlying goal. Open rate tells you whether the subject line earned a click. It doesn’t tell you whether the email earned attention from the right person. Reply rate counts responses, but “not interested” and “tell me more” both increment the same counter. Meeting booked tells you someone agreed to spend 30 minutes, not that they’re a qualified buyer who converted because the copy was calibrated to their decision-making style.

When a team optimizes against a proxy, they tend to close the gap between their copy and the metric. They do not tend to close the gap between the metric and the underlying goal. If anything, sustained proxy optimization often widens that second gap. The copy gets better at producing the metric. The metric gets worse at measuring the goal.

That’s Goodhart’s trap closing. The measure was accurate when you started. Optimization made it inaccurate. You didn’t notice, because you were watching the metric, not the thing the metric was supposed to represent.


Three patterns where it plays out in B2B copy

The curiosity-bait subject line spiral

A team A/B tests subject lines. The winning format is vague and slightly provocative: “Most marketing teams miss this,” or “The reason your open rates are lying to you.” Opens go up. Clicks go up. Replies go down.

What happened: curiosity-bait subject lines optimize for a universal human behavior. People click on things that create mild information gaps. That behavior is real and measurable. The problem is it’s not selective. Curiosity-bait opens attract a broad, undifferentiated pool of readers. The email body, written for a specific buyer with a specific problem, now has to earn the attention of everyone who clicked, including many who are not the buyer you’re selling to.

The next test is another subject line test. The curiosity-bait format keeps winning on opens. Nobody questions whether open rate is still the right test variable, because the spreadsheet looks good.

Three months later, reply rate is down 30%. The copy has been optimized toward a signal that correlates weakly with qualified interest and negatively with buyer-specific resonance.

The engagement score drift

A marketing team uses a copy scoring tool that rewards engagement signals: high-affect language, action verbs, short punchy sentences, forward momentum in the copy structure. Their content scores keep improving quarter over quarter. Their deal velocity in the CFO segment slows.

What happened: the engagement score was calibrated against general engagement research, broad behavioral patterns that predict response across mixed audiences. It was not calibrated against the specific information-processing style of a high-Conscientiousness buyer, which is how CFOs tend to read.

High-Conscientiousness readers process evidence before affect. They weight specificity over urgency. They disengage from copy that leads with emotional pull before establishing a verifiable claim. The engagement score was optimizing the copy toward patterns that work on high-affect readers and away from patterns that work on evidence-first readers. The CFO segment was not in the training data.

The performance pivot trap

An email test produces a clear winner: the version with more specificity and fewer warm-up sentences outperforms the control. The team abstracts the learning: “be more specific.” They apply it across all copy variants for all segments.

Results are mixed. Open rates are flat. Some segments improve. One segment gets notably worse.

What happened: specificity improved performance in the original test segment because those buyers were high-Conscientiousness readers who reward evidence density. Specificity isn’t a universal signal. For a different segment (high-Openness decision-makers in the early stages of a buying consideration), it can read as over-prescribed and under-interesting. High-O readers want conceptual range, a frame they can extend, an idea they haven’t seen yet. A barrage of specifics can feel like a product spec sheet when they were expecting a new angle.

“Be specific” was the right learning for one personality profile and the wrong generalization across all of them.


What actually breaks when you optimize the score

The damage is invisible in the metrics, which is what makes Goodhart’s trap so durable.

No individual test reveals that your copy is losing personality alignment with a specific buyer segment. Each test surfaces the winning variant on the metric you’re optimizing. What’s invisible is the cumulative drift across 20 tests, where the copy has migrated away from the configuration that was producing qualified interest at baseline.

The thing that breaks is specificity of fit. Copy that performed because it was calibrated to a particular buyer’s information-processing style gets re-optimized toward a proxy metric that doesn’t know that buyer exists. The metric goes up. The underlying resonance goes down. The score and the goal have been decoupled.

The team doesn’t see this because they’re looking at the metric. The metric doesn’t show what it’s lost track of. Every number in the dashboard is technically correct. The interpretation has become wrong.

This is Goodhart’s Law in full operation. You can’t see the corruption from inside the measurement system, because the measurement system is what’s corrupted.


What personality-grounded scoring changes

COS scores copy against the OCEAN personality profile of the target buyer. OCEAN is the Big Five personality model (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism), grounded in 860+ peer-reviewed papers on how personality traits shape information processing, claim evaluation, and decision-making.

That’s a different measurement surface than engagement metrics. COS isn’t measuring whether the copy is stimulating in the abstract. It’s measuring whether the copy matches the cognitive style and emotional register of the specific buyer role you’re targeting.

A high-Conscientiousness CFO needs evidence-dense, prevention-framed copy with verifiable claims and low emotional load. A high-Openness VP Marketing needs conceptual breadth, possibility framing, and an idea she can extend. Copy that scores well on OCEAN alignment for one of those buyers and poorly for the other isn’t a failure. It’s a targeting decision made visible.

This is what personality-grounded measurement prevents: the Goodhart trap. When you optimize for OCEAN alignment with a specific buyer, the score is grounded in a stable underlying structure: the buyer’s actual cognitive preferences, which don’t shift because you ran 20 email tests. The proxy hasn’t become detached from the goal, because the proxy is measuring something closer to the goal directly.

It won’t tell you that your open rates are up. It will tell you whether the copy that produced those open rates is actually calibrated for the buyer you’re trying to convert. Those are different questions. The second one is the one that matters when open rates are up and reply rates are down.


Before you run the next test

There are three questions worth asking before you iterate on the next variant.

What is the score actually measuring? Not in theory, but given the optimization pressure that’s been applied to it over the last six months. Open rate was supposed to measure “did this message connect with a qualified buyer enough to read further.” Does it still do that, or has your copy been tuned toward behaviors that produce opens without qualifying readers?

Is the metric still correlated with the thing you care about? The correlation was probably real when you started. Sustained optimization pressure tends to erode it. The empirical check is to look at a metric that isn’t in your optimization loop. If reply rates or meeting-booked rates are moving in the opposite direction from open rates, the correlation has broken. Goodhart has closed the trap.

What’s your independent check? The test data tells you how performance changed on the metric. It doesn’t tell you whether the metric still measures performance on the underlying goal. You need a signal that isn’t inside the optimization loop. For copy, that’s personality-fit scoring: it evaluates the copy against the buyer’s decision-making structure, not against the behavioral patterns that happened to produce clicks.

The measurement trap is not a failure of rigor. It’s a failure of scope. The discipline is real; the question is whether it’s measuring the thing that determines whether the copy works for the right buyer.

Metrics are not compasses. They’re instruments with a known limitation: sustained optimization pressure eventually breaks the reading. The question is when you last checked whether the needle still points north.


Start with the buyer, not the metric

The practical fix is not to stop testing. It’s to add a measurement layer that doesn’t break when you optimize against it.

COS generates personality-grounded B2B copy and scores existing drafts against four frameworks: Engagement, Personality fit, Strategic Clarity, and Framing Strategy. The Personality fit score tells you how well the copy aligns with the OCEAN profile of the target buyer role: not how engaging it is in the abstract, but how calibrated it is to how that specific reader processes information and evaluates claims.

Run your current CFO-segment cold outbound through it. Run the version your test identified as the winner. Check whether the winner is more personality-calibrated for that buyer, or just better at producing the behavioral signal your test was measuring.

If those two things are aligned, you’re in good shape. If they’re not, you’ve found where the measurement trap closed.

For a related pattern in the B2B sales context, where metric-goal gaps create stalls and delays at procurement: The Real Reason Demo Calls Stall at Procurement.

Know what the score is measuring. Know whether it’s still measuring that. The data inside an optimization loop is the last place to look for that answer.