Skip links

What a Fine-Tuned Model Cannot Do

Why this article exists

Most writing about fine-tuning covers what it achieves.

Knowing the limits matters more, because that is where people get hurt — relying on output that looked right and was not.

The best model cards state their own limitations plainly. AboutKnowledge’s two published quality models both say, in their own documentation, that they should not replace human auditors or regulatory review, and that compliance decisions must be verified against the official standard text.

That is what a professionally published model looks like. Its absence elsewhere is a warning.

1. It cannot be relied on to be correct

A fine-tuned model is more reliable on its subject. It is not reliable.

It produces the most likely output given its training. On typical cases that is very good. On unusual ones it can be confidently wrong.

And there is no signal. A wrong answer reads exactly as well as a right one — same tone, same structure, same apparent confidence.

Which means every use needs a review step, sized to the consequences. Drafting a checklist? A glance. Deciding compliance? A qualified person.

2. It cannot stay current

A model knows what it was trained on, at the time it was trained.

Standards get revised. Regulations change. New requirements are added.

The model does not know. It will answer with what it learned, confidently, about a version that has been superseded.

Two implications:

Retraining is a recurring cost, not a one-time build.

For anything where the current version matters, verify against the source. A model is a way to find your way around a standard, not a substitute for reading it.

FIGURE 1: THE LIMITS THAT MATTER MOST

Not reliably correct

  • Confidently wrong on unusual cases, with no signal.

Cannot stay current

  • It knows what it was trained on. Standards get revised.

Cannot explain its reasoning

  • It produces an answer, not an auditable justification.

Only as good as its training data

  • Errors in the source are learned thoroughly.

3. It cannot explain itself in a way you can audit

It produces an answer. It can produce text that looks like reasoning.

Those are not the same thing.

The explanation is generated alongside the verdict, not derived from a process you can inspect. It usually matches. It is not a proof.

Which matters when somebody asks how a conclusion was reached.

For an operational nudge, fine. For anything a regulator, auditor or certification body may examine, you need a documented method a person can defend — not a model output.

4. It cannot exceed its training data

Errors in the source get learned thoroughly.

Both AboutKnowledge model cards say this directly: synthetic training data can contain occasional grammar artifacts, or mix standards in generic answers.

That is honest, and it is a general truth about fine-tuning. A model trained on a knowledge base inherits that knowledge base’s mistakes — and reproduces them with the same confidence as everything else.

Nor can it answer what it was never shown. Ask about a standard outside its training and you get a plausible answer assembled from adjacent knowledge. That is worse than an admission of ignorance, because it looks like an answer.

5. It cannot handle what it was not built for

A model fine-tuned on one subject is unremarkable outside it.

That is the trade you made. All the capacity on one thing means less of everything else.

Practically: do not use a quality-management model for general questions. It will answer, and the answer will be worse than a general model’s.

Know what your model is for, and use it for that.

6. It cannot make a decision with consequences

The line that should not move.

Compliance decisions. Regulatory submissions. Anything a certification body reviews. Anything where money moves.

A model can prepare, triage, draft and flag. All of that saves real time.

A person decides.

Both of these published models state exactly this position — useful for drafting and triage, not a replacement for human auditors or regulatory review. Written by the people who built them, which is the most credible place for that caveat to come from.

FIGURE 2: WHAT IT DOES AND DOES NOT DO

Genuinely useful for

  • Drafting a first version
  • Triage — what to look at first
  • Consistency across assessors
  • Finding your way around a subject

Not for

  • The decision itself
  • Anything a regulator will examine
  • Quoting source wording as authoritative
  • Replacing qualified judgement

7. It cannot tell you when it is unsure

Some models can be prompted to express uncertainty. It is not reliable.

A model does not have a calibrated sense of its own confidence. It will produce an answer with the same apparent certainty whether the case is typical or unlike anything it has seen.

Which means you cannot use its confidence as a filter. “Only review the ones it was unsure about” does not work.

The practical consequence: review a sample of everything, not just the cases that looked doubtful.

What to check before relying on one

Six questions, for any fine-tuned model — yours or somebody else’s.

What was it trained on? Real data, synthetic, or a mix. From what source.

How much? Thousands or hundreds of thousands. Volume tells you something.

What does the model card say about limitations? If it says nothing, be careful. Every model has limits; a card that lists none is incomplete rather than exceptional.

What is the base model and licence? An adapter inherits its base model’s terms.

When was it trained? For subjects that change, age matters.

Where does a person review? If the answer is nowhere, that is the problem to fix first.

FIGURE 3: WHAT TO LOOK FOR IN A MODEL CARD

Stated limitations

  • Every model has them. A card listing none is incomplete.

Training data described

  • Real or synthetic, from what source, how much.

Base model and licence

  • An adapter inherits the base model’s terms.

An intended use, stated

  • What it is for, and what it is not for.

The realistic position

A fine-tuned model is a capable assistant on its subject, whose work you check.

That is a smaller claim than the marketing around AI, and a bigger one than the sceptics allow.

For drafting, triage and finding your way around a body of knowledge, it saves real time.

For decisions with consequences, it prepares the ground and a person decides.

That position is not caution for its own sake. It is what the people who build these models say about their own work — and they are in the best position to know.

The short version

A fine-tuned model cannot be relied on to be correct, cannot stay current, cannot explain itself in an auditable way, and cannot exceed its training data.

It also cannot tell you when it is unsure, which is why sampling everything beats reviewing only what looked doubtful.

Use it to draft, triage and orient. Let a person decide anything with consequences.

And when evaluating any published model, read the limitations section first. If there isn’t one, that tells you something.

Considering a fine-tuned model for a business task?

Get in touch. We will be specific about where it helps and where it should not be trusted — that conversation is worth having before the build, not after.

Leave a comment

Drag