How to test an AI feature across different kinds of users (and why you rewrite rules instead of patching them)
The way to test an AI feature across different kinds of users is to build a small internal tester that runs the same scenario through several real user profiles, read the outputs side by side, and rate each one. When the same miss shows up across several runs, go back and rewrite the instruction that caused it rather than adding another "don't do this" line to the end of the prompt. Then rerun everything to confirm the fix didn't break a different kind of user. That's the loop we used to ship AI trip-prep emails to tens of thousands of travelers at Pangea.
Why you need it: same product, very different right answers. A weekender doing four days in Lisbon, a nomad who anchors in one city for months, and a family in Japan during Golden Week all get the same feature. What a good email looks like is completely different for each. A prompt that nails one can quietly fail the other two, and you won't see it by testing one trip at a time.
The tester. We built a simple internal app: pick a real traveler profile, a destination, and dates, and it generates every email in the sequence. I spent a lot of time in it, running the same trip through different profiles and reading the outputs next to each other. You don't need to be an engineer to build or use something like this. It's a form and a results page.
Rate every run, and write down what it should have said. A score alone doesn't tell you what to fix. The note ("should have flagged that this city doesn't use Uber") is what you'll actually use later.
Look for the pattern, then fix the root rule. One bad output is an anecdote. The same miss across several runs is a flaw in the original instructions. The tempting fix is a new line at the bottom of the prompt. Do that enough times and the prompt ends up scarred over with patches that contradict each other. Instead, find the rule that produced the miss and rewrite it.
Rerun the same trips. After every rewrite, run the same set of profiles and trips through the tester again. You're checking two things: that the miss is gone, and that the fix didn't break the emails for a different kind of traveler.
The same approach works for any GTM agent. If your outbound agent writes to a CFO at a 50-person startup and a VP of Sales at an enterprise, those are your weekender and your family. Build the tester before you build the fifth patch.
Source: originally published on The Agent GTM Newsletter.