# The Evaluation Gap ## The company gave no proof at the level of the model for the AI claims The company gave no model cards, no descriptions of training data, no test sets, no precision figures, no latency numbers and no validation methods. This applies to each model, and four facts are unknown. The first is the foundation model behind the domain-specific language model. The second is the fine-tuning method, which is parameter-efficient tuning or continued pre-training. The third is the quantity of data in the dataset. The fourth is the method to label the data. > [!warning] Largest missing proof > All quantified outcomes are **self-reported and unattributed, and have no stated baseline or method**. The outcomes are an increase in efficiency, an annual saving, and a decrease in mean time to repair. A ramp-time improvement figure in a customer letter is a *target* in a proof-of-concept. It is not a result. ## Three layers, three answers | Layer | Probably | Why | | --- | --- | --- | | Language-model agents | Value comes more from retrieval, prompting and workflow architecture than from training | Across enterprise deployments, retrieval quality usually has the largest effect. The fine-tune must be better than an off-the-shelf open-weight model with the same retrieval, on the tasks of the company. | | Classical ML (root-cause analysis, anomaly detection, virtual metrology, optimization) | **Actual ML** | The methods are standard. But it is scarce expertise to apply them correctly to factory data. | | Document-generation agents (8D reports, compliance reports, forecasts) | Templated drafting on retrieved data | They are useful, they have low technical depth, and a competitor can copy them easily. | > [!important] The largest contract can be external to the documented architecture > Defect detection for footwear and fabric is a computer-vision problem. The architecture diagram has no vision model, no labeling pipeline and no integration with inspection hardware. ## The test that resolves the question A third party runs a **blind test** on a held-out set of domain tasks. The test compares the fine-tuned model with an off-the-shelf open-weight model. The two models use the same retrieval pipeline. The test also has more parts. The parts are regression sets for agents and hallucination tests. The parts are also dashboards that show changes in model behavior in live deployments, and logs of the human-approval step. ## Why it matters The selling proposition of the company is "Decision-grade output". Without proof from tests, a reader cannot find a difference between the company and a correctly built retrieval chatbot. [[Evals]], [[Domain Experts as Eval Builders]] and [[Where Domain Evals Matter Most]] give the treatment of this topic in the vault. [[Evidence Hierarchy]] is the grading frame. ## Related - [[Foreman MOC]] - [[Claims Versus Evidence]] - [[Human-in-the-Loop Systems]] - [[The Tribal Knowledge Thesis]]