Generative AI consulting, as we practice it, is an engineering engagement, not an inspiration session. The questions that decide whether your system works are unglamorous: which model, chosen by measured performance on your data rather than by brand; what the model is allowed to read and to touch; what a correct answer is and how you would know; and what each call costs once the volume is real. Get those four wrong and no amount of prompt polish recovers it.
The failure we are built to prevent has a familiar shape. A pilot demos beautifully — clean inputs, a friendly audience, ten hand-picked examples — and then meets production: the scanned PDF at an angle, the customer who writes in two languages, the record that contradicts the CRM, the question the documents do not answer. The pilot was never tested against any of that, so nobody finds out until it is live and wrong in front of a customer.
So we invert the order. Before rollout, we define what “it works” means — pass and fail, written down, measured on a sample of your real data including the ugly cases. A system that has not been evaluated on your data has not been evaluated; it has been rehearsed.
If a generative model is the wrong tool for your problem, we will say so on the call. A decent share of what gets pitched as a GenAI project is a lookup, a rules engine, or a form redesign — cheaper, faster, and it never hallucinates.