Free: the GTM BlueprintThe stack by funding stage, the CRM data model written down, and a 90-day plan you can run on Monday.Get the blueprint
RevOpsXL
GTM Diagnostic Ask about the audit

Teach the model your deals, not the internet

A generic model knows the internet; it has never seen how your deals close, and that gap is the whole game.

In short

A general model knows the world at large and your business not at all. Ground it in your real, unsummarised deals, emails and objections - retrieval over your own records, not summaries - and it starts to reason the way your best rep does.

On this page

A general-purpose model has read most of the internet and none of your deals.

It knows what a SaaS contract is in the abstract. It has never watched one of yours stall for six weeks because procurement went quiet, then close the day the champion got promoted. It doesn't know that in your market the second call is where deals die, or that the objection about onboarding is really about budget.

That knowledge isn't on the internet. It's in your CRM, your email, your call notes. The whole task is getting it out of there and in front of the model.

Everyone is shipping the same base model. The difference comes from what you feed it.

What the internet taught the model, and what it didn't

The base model is general knowledge. Broad and useful, but completely average, because it was trained on the average of everything.

Your edge is the specific, hard-won stuff: which signals predict a close in your market, how your buyers talk, the objection that sounds fatal but never is. A generic model pulls everything toward the mean of the internet. Your best rep does the opposite. They have seen your deals, not everyone's.

The internet made the model articulate without making it yours. An articulate model that doesn't know your market hands you a confident answer that is true in general and wrong for you - the most dangerous kind of wrong, because it reads so well.

So the question isn't which model is smartest, but which one has seen your deals.

Retrieval over your own records beats a cleverer model

The way to close that gap is to give an ordinary model access to your own records at the moment it answers, rather than reach for a bigger one.

In plain terms: asked to judge a deal, the model first pulls the most relevant real examples from your own history - similar deals, how they went, the emails that turned them - and reasons over those. That's retrieval, and it's the gap between a model guessing from the internet and a model reasoning from your evidence.

The knowledge lives in your data, not in the model's weights. That's the shift most people miss. You're not making the model smarter; you're making the evidence you already have reachable. It's the core of how I build these systems.

Isn't this just fine-tuning?

No, and the difference matters if you're paying for it. Fine-tuning adjusts the model's weights - good for teaching it a style or a format, bad for teaching it facts that change every week.

Your pipeline changes daily. Deals close, objections shift, a new competitor turns up in the notes. Retrain the model every time and you'll go broke and still be a week behind. Retrieval reads the current records at the moment of the question, so the model always works from today's evidence, not a snapshot from the last training run. For deal knowledge, that's what matters: you want fresh evidence, not something baked in from the last run.

Real and unsummarised beats clean and summarised

The instinct to fight is the urge to tidy the data first.

Summarising your deals before you feed them in feels responsible, but it does the opposite of what you want. A summary keeps the conclusion and throws away the reasoning - and the reasoning is the only part worth having. 'Deal lost to budget' is a summary. The actual thread where the buyer cooled from keen to gone over eleven days, that's the lesson. Feed the model the summary and you've taught it your labels. Feed it the whole thread and it learns the pattern instead.

The mess is the signal. The hesitations, the exact wording, the six-week silence - that's the texture a model needs to reason like someone who was in the room.

Synthetic data teaches it to sound right, not be right

There's a shortcut on offer: generate synthetic deals, or have a model write example objections, and train on those. Don't.

Synthetic data is a model's idea of what your deals look like. Train on it and you get a model imitating itself - fluent and plausible, but hollow. It'll produce a beautiful rationale for why a deal will close, about a deal that never existed. You've taught it to sound right, which is the exact failure you were trying to avoid.

Real data is harder to work with. It's also the only kind that carries the thing you're paying for: what happened.

What 'reasons like your best rep' actually means

Your best rep isn't smarter than the model. They have seen a few hundred of your deals and remember how they went.

Ground a model in the same evidence and you get something close to that instinct - available to everyone, at three in the morning, on every deal in the pipeline. It reads a new opportunity, recalls the twelve most similar deals you've run, and tells you what usually happens next, with the real examples attached so you can check its work. This is what I mean when I write about training agents on real data.

A generic model knows what a deal is. Only your data knows how yours close.

Teach it the second thing.

Common questions

Should I fine-tune a model on my sales data?

Usually not for deal knowledge. Fine-tuning is good for style and format but bad for facts that change weekly; retrieval over your live records keeps the model working from today's deals without constant retraining.

Why not use synthetic data to train a sales AI?

Synthetic data is a model's guess at what your deals look like, so training on it teaches the model to imitate itself - fluent but hollow. Only real, unsummarised deals carry what actually happened, which is the part worth learning.

What does it mean to ground a model in your own data?

It means the model pulls relevant real examples from your CRM, email, and notes at the moment it answers, and reasons over those instead of over generic internet knowledge. That's how it starts to sound like someone who has seen your deals.

Where does yours break?

Send over the detail and we’ll tell you what we would check first. A senior expert reads these, not a bot, and you’ll get a straight answer either way.

Rather look around first? Run the free GTM diagnostic, or follow RevOps XL on LinkedIn for a new field note twice a month.

← More field notes