Determinism before LLMs: when a dictionary or a well-written function beats the model
Aduanera del Estrecho is a customs brokerage based in Algeciras, in southern Spain. It has 42 employees and turns over around €8 million a year clearing containers through the port. In February 2026 the technical director hires a consultancy to "add AI" to the tariff classification workflow. The consultancy wires GPT-4o into the ERP input: every supplier invoice line goes through the model to be assigned its 10-digit TARIC code (TARIC is the EU's integrated customs tariff). The system processes around 18,000 lines a month. The OpenAI bill for the first month: €1,240. The error rate: 11%. A senior customs broker corrects every line by hand, because a wrong TARIC code means between €300 and €2,000 of duty settled incorrectly.
When I come in to review the workflow, the question isn't how to improve the prompt. The question is why there's an LLM there at all. TARIC is a closed catalogue. The European Commission publishes it, you can download it as XML, and it changes twice a year. We replace the LLM with a matching function against the official catalogue plus a synonyms table that the senior broker maintains. Monthly cost: €0 in API fees. Error rate: 0.7%, all of them ambiguous cases that the catalogue itself flags as such.
The model wasn't failing because it was bad. It was failing because it shouldn't have been there.
This is the expensive mistake I see repeated in 90% of legacy architectures with AI: a model applied to a task that code or a well-maintained catalogue already solves.
And it applies today.
The five families of tasks that don't need a model
Before you bring in an LLM, rule out five types of task where the model always loses on cost, accuracy or traceability. If your problem fits any of these five families, you don't need the LLM.
- Classification against a closed catalogue. TARIC codes, CNAE codes (Spain's classification of economic activities), CPV codes (the EU public procurement vocabulary), IBAN formats by country, postcodes, province by prefix. The set of valid answers is finite and published by an official body. The correct answer is either in the table or it doesn't exist.
- Format validation with known rules. Spanish DNI and NIE identity numbers, IBAN, CIF company tax IDs, number plates, Social Security numbers, cadastral references. Each one has a documented check-digit algorithm. A 20-line function returns true or false with no ambiguity.
- Deterministic transformation from one value to another. ISO date to DD/MM/YYYY, upper case to lower case, stripping accents for search, calculating age from date of birth, converting currency at the day's European Central Bank (ECB) rate. One to one, no judgement involved.
- Routing based on explicit rules. If the amount exceeds €3,000, the finance director approves. If the supplier is on the blacklist, block. If the country isn't on the whitelist, escalate. A decision tree does it in a millisecond with perfect logs.
- Exact or near-exact search in your own database. Finding a customer by email, a product by SKU, a contract by case number. A database index resolves it in microseconds with 100% accuracy.
These five families cover a large share of the task volume in an average business workflow. Putting an LLM on any of them means paying more to do worse something you already knew how to do well.
Rule: if you can write the decision criterion in a sentence that starts with "if the value is in this list" or "if it matches this pattern", you don't need the LLM.
How much it costs to use a model where it doesn't belong
On a deterministic task, the API bill is only part of what an LLM costs. The real cost is the sum of four layers that almost nobody counts when they buy the consultant's proposal.
The figures below come from a comparison I ran on Aduanera del Estrecho's real workflow during the first quarter of 2026, using OpenAI's and Anthropic's public prices in force in March 2026 (GPT-4o at $2.50 per million input tokens, Claude Sonnet 4.5 at $3 per million). Treat them as an order of magnitude, not a universal benchmark.
Cost per 10,000 TARIC classification lines, compared:
- LLM without catalogue. €68 in API fees, 11% error, 6 hours a week of a broker reviewing and correcting. Total cost with the broker's hour at €32 gross: €260 per 10,000 lines.
- LLM with the catalogue loaded into the prompt. €214 in API fees because the TARIC catalogue takes up a lot of tokens, 3% error, 2 hours of review. Total cost: €278. Worse than without the catalogue, because the context tokens sent the bill through the roof.
- LLM with RAG against the catalogue. €41 in API fees, 2.5% error, 1.5 hours of review. Total cost: €89. Better, but with extra infrastructure (vector database, ingestion pipeline, monitoring).
- Deterministic function against the official catalogue. €0 in API fees, 0.7% error, 20 minutes a week reviewing cases that the catalogue itself flags as ambiguous. Total cost: €11.
The deterministic function is 24 times cheaper than the LLM without catalogue and 8 times cheaper than the LLM with RAG. And its accuracy is higher in all three comparisons.
In plain terms: when the task is deterministic, every euro you spend on the LLM buys you a worse result.
There's one more layer that rarely makes it into the accounts: traceability. When the broker has to justify to a customs inspector why a given TARIC code was assigned, with the deterministic function they show the catalogue line that matched. With the LLM they show a prompt and an answer that the model might generate differently next time. That cost doesn't appear on the bill until the inspection arrives.
Why the model makes up answers when the criterion already exists
This is the part many people find hard to accept. An LLM isn't an information retrieval system. It's a probabilistic generation system. When you ask it to assign a TARIC code and it hasn't seen that exact product in training, it doesn't say "I don't know". It generates the code that is statistically most likely given the context.
And the most likely answer isn't the correct answer.
A real case from Aduanera del Estrecho's workflow: a Chinese supplier describes a product as "aluminium foil roll, food grade, 30 micron". The correct TARIC code is 7607111990: aluminium foil less than 0.021 mm thick, for food use, unbacked. GPT-4o consistently returns 7607192090, which is backed aluminium. The duty difference is 4.2 percentage points on the container's CIF value. On a €24,000 container, that's €1,008 settled incorrectly. Multiply that by the roughly 40 containers a month from this supplier and you get €40,000 a month that customs will claim back at the next inspection.
The model isn't broken. It's doing exactly what it was asked to do: generate the most likely answer. The problem is that the most likely answer doesn't match the correct one when the decision criterion is already codified in an official document.
The LLM doesn't consult the European Commission. It consults its statistical memory of having seen lots of texts about aluminium.
This gets worse in three scenarios where you should never let a generative model decide:
- When the answer has legal or tax consequences. TARIC, VAT by category, IRPF withholdings (Spanish personal income tax), Social Security codes. The regulator expects traceability. The model doesn't have it.
- When the catalogue changes and the model doesn't know. TARIC is updated twice a year. A model trained in January 2025 doesn't know the codes added in July. Asking it guarantees you outdated answers.
- When the same input must always produce the same output. An LLM with a temperature above zero can give different answers to the same question. A catalogue can't. If the auditor asks you to reproduce a classification from three months ago, the deterministic function gives it to you. The LLM doesn't.
Symptom: the model returns plausible answers that vary from run to run. Fix: take it out of the workflow and replace it with a lookup against the official source.
The living catalogue pattern, maintained by the business
When I replace an LLM with a deterministic function, hardcoding a table in the code isn't enough. That's what burns out the technical team within three months, when the catalogue changes and nobody knows how to update it. The pattern that works is what I call a living catalogue, and it has four layers.
The first layer is the authoritative source. The European Commission publishes TARIC. INE, Spain's national statistics institute, publishes CNAE. Correos, the Spanish postal service, publishes postcodes. SWIFT publishes IBAN formats by country. Each catalogue has an external owner who decides which codes exist. Download the original catalogue in its official format (XML, CSV, JSON) and store it untouched in a versioned directory. Never edit this file.
The second layer is the derived catalogue: the authoritative source transformed into the format your system needs, whether that's a Postgres table, a JSON file in an S3 bucket or a pickle file in memory. It regenerates automatically every time a new version of the official catalogue arrives. Zero manual intervention.
The third layer is the synonyms and overrides table. This is where the business is in charge. At Aduanera del Estrecho, the senior broker maintains a spreadsheet with three columns: the supplier's free-text description, the official TARIC code that applies, and notes. When a Chinese supplier writes "aluminium foil roll, food grade, 30 micron", the table has a line that says "aluminium foil food grade < 40 micron -> 7607111990". The matching function checks the synonyms table first and only falls back to the official catalogue if there's no match. The table is maintained by the business person who knows the domain, not by the engineer.
The fourth layer is the log of unresolved cases. When the function finds no match in either the synonyms table or the official catalogue, the case goes into a review queue. The senior broker goes through it once a week, decides which code applies and adds the corresponding line to the synonyms table. The system learns, but a person with judgement directs the learning, not a model.
This pattern has three advantages no LLM gives you:
- Auditable. Every classification decision leaves a trail: which catalogue or synonyms line matched, when it was added, who added it. Tax inspectors appreciate that.
- Deterministic. The same input always produces the same output. If something changes, it's because someone changed it on purpose, and that's on record.
- Cheap. No API. A small server handles millions of queries a day. The infrastructure is maintained by the same person who already looked after the ERP.
At Aduanera del Estrecho the synonyms table grew from 0 to 340 lines in the first two months. From the fourth month, growth dropped to 3-4 new lines a week. The system stabilised because suppliers keep sending the same products.
An LLM never stabilises. It charges you the same every month because it doesn't learn from your corrections unless you set up fine-tuning, and setting up fine-tuning to classify TARIC codes is exactly the kind of architecture you need to avoid.
When the dictionary stops being enough
I'm not saying the LLM is useless. I'm saying it's useful for specific tasks where the variability of the input doesn't fit in any catalogue, however large. There's a clear line that marks when the dictionary falls short and you need to move up a level.
Move up a level when any of these four signals appears:
- The input is free text with variable semantic structure. A customer complaint written in an email, a description of medical symptoms, an open answer to a survey. There's no catalogue of "possible complaints". The model adds value because it brings understanding.
- The answer requires combining several sources of knowledge. Writing an executive summary from three different reports, putting together a legal answer that draws on case law and regulation, generating a call script that uses the customer's history. Here the model composes.
- The task is inherently subjective. Writing an email in a professional tone, suggesting a catchy headline for a blog, translating while keeping the cultural register. There's no single correct answer. The model generates plausible variants.
- The cost of an error is low, or human review before publishing is mandatory. A draft reply to a customer that a person reviews, a code suggestion a programmer accepts or rejects, a meeting summary shared for validation. The model assists; it doesn't decide.
When none of these signals appears, the LLM is surplus. When one does, the LLM can add value. Even then, the next question is which part of the workflow can stay deterministic.
A concrete example. Consultoría Administrativo, an accounting and tax advisory firm in Seville with 18 employees, wants to classify and answer incoming client tickets. Workflow analysis:
- Classifying the ticket by area (tax, employment, accounting, corporate): deterministic task. Done with a keyword list the firm itself maintains. 92% accuracy without an LLM.
- Detecting urgency: deterministic task. If the ticket mentions "formal demand", "notice", "deadline", "today" or "tomorrow", it's flagged as urgent. A simple regex.
- Routing to the responsible adviser: deterministic task. A client-adviser table maintained in the CRM.
- Drafting a reply in the firm's tone: here, yes, an LLM. The input is free text and the output is free text with professional nuance.
The workflow goes from "LLM across the whole chain" to "LLM only on the last mile". API cost drops by 78%. Latency drops by 60%. Routing accuracy rises from 71% to 94%. The LLM is still useful, but only at the step where it adds something.
Rule: the LLM goes into the workflow where the criterion can't be written down. If it can be written down, write the criterion.
How to spot legacy code with a misplaced LLM
When I review a legacy architecture, I have six quick questions that tell me within 20 minutes whether the LLM is in the right place. They're questions you can ask yourself before hiring anyone.
- What is the set of possible answers? If it's finite and can be published as a list, you don't need a generative model. A catalogue is enough.
- Does the correct answer change over time for technical or regulatory reasons? If so, an LLM trained at a fixed date will give you outdated answers. You need an authoritative source you can query in real time, with or without a model on top.
- Must the same input always produce the same output? If so, an LLM with a temperature above zero is ruled out. If you're willing to set the temperature to 0 and still accept variability between model versions, check whether the task is deterministic.
- Can a person write the decision criterion in one sentence? If so, that criterion is code. An if, a lookup, a regex. The LLM is surplus.
- Is there a downloadable official data source that already contains the answer? TARIC, CNAE, BOE (Spain's Official State Gazette), the Official Journal of the EU, the medicines catalogue of AEMPS (the Spanish medicines agency), Correos postcodes. If the answer is in a public file, querying the file is cheaper and more accurate than asking the model.
- Is the cost of an error high? If so, an LLM without traceability is a liability. Either you add a deterministic validation layer afterwards or you take the model out of the critical path.
If you answer "yes, closed catalogue" or "yes, there's an official source" to any of these six questions, you have a misplaced LLM in the workflow.
Common mistakes when redesigning a workflow with determinism
Done well, a redesign like this takes two or three weeks. Done badly, it drags on for months and stays half-finished. These are the five mistakes I see again and again.
- Mistake 1: Replacing without measuring the starting point. Symptom: you remove the LLM and can't show how much accuracy improved. Fix: before touching anything, run 500 cases manually, measure accuracy and cost, and save the result. Without a baseline there's no improvement you can prove.
- Mistake 2: Hardcoding the catalogue in the code. Symptom: six months later nobody knows how to update the TARIC codes because they sit in an array inside a
.pyfile that only the original engineer understood. Fix: the catalogue lives in a table or in a file separate from the code, with a documented process for updating it from the official source. - Mistake 3: Not giving the business ownership of the synonyms table. Symptom: the domain expert has to open a ticket with the technical team every time they want to add a line. Within two months the table is out of date. Fix: the synonyms table lives in a spreadsheet or an admin panel the domain expert can edit without technical help.
- Mistake 4: Keeping the LLM as plan B with no exit criterion. Symptom: when the catalogue gives no match, the workflow calls the LLM as a fallback. With no record of those cases, the LLM goes back to deciding on its own and traceability is lost again. Fix: when the catalogue gives no match, the case goes to a human review queue. The LLM is only used if the person explicitly authorises it for that type of case.
- Mistake 5: Not versioning the catalogue. Symptom: in January you classified 12,000 products with a TARIC code. In April a client complains and you need to reproduce January's classification. You can't, because the catalogue was updated in March and you didn't keep the previous version. Fix: store every version of the official catalogue with its date. The classification function can take a date parameter to query the version in force on that day.
None of these mistakes is technical. All five are process mistakes. That's why the redesign has to be led by someone who understands data, business and architecture at the same time. A pure backend developer stops at "I built the function". A pure consultant stops at "I defined the process". The bridge between the two is missing.
Frequently asked questions
When shouldn't you use an LLM in a business workflow?
When the task fits one of these five families: classification against a closed catalogue, format validation with known rules, deterministic transformation from one value to another, routing based on explicit rules, or exact search in your own database. In all these cases the LLM costs more, is less accurate and is harder to audit than a deterministic function. The quick rule: if you can write the decision criterion in a sentence that starts with "if the value is in this list" or "if it matches this pattern", you don't need the LLM.
Which tasks are better solved with code or a dictionary than with AI?
Any task with a single, verifiable answer. Assigning a CNAE code to a company from its main activity (5,000 possible codes published by INE, Spain's national statistics institute), validating a Spanish IBAN (known check-digit algorithm), calculating the VAT that applies to a product from its family code (a table maintained by the tax department), routing an email to the right department by sender (an account table maintained by HR), converting currencies at the day's ECB rate (the European Central Bank's public endpoint). None of these tasks gets better with an LLM.
How do you know whether a task needs a generative model or determinism is enough?
Ask yourself four questions. One: is the set of possible answers finite and publishable as a list? If so, you don't need a model. Two: must the same input always produce the same output? If so, you don't need a model with variability. Three: can a domain expert write the decision criterion in one sentence? If so, that criterion is code. Four: is there a downloadable official source with the answer? If so, querying the source is cheaper and more accurate. If you answer "no" to all four and the task involves free text with variable semantic structure, then yes, the model adds value. Otherwise, go deterministic.
Why can a code catalogue be better than asking ChatGPT?
For three reasons the bill doesn't show. First, accuracy: the official catalogue is the authoritative source, while the model is a statistical approximation that can hallucinate plausible but wrong codes. Second, traceability: the catalogue returns the exact line that matched, which you can show at an inspection; the model returns an answer that can vary between runs. Third, marginal cost: each catalogue query is practically free, while each model query has a per-token price that multiplies with volume. In tariff, accounting or regulatory classification, the catalogue wins on all three counts.
What do I do if my current workflow already has a misplaced LLM?
Work in this order. One: measure the current state with 500 manual cases. Accuracy, cost per 1,000 cases, latency. Without a baseline you have nothing to negotiate with, either with the supplier or with management. Two: identify the step in the workflow where the model sits and apply the six detection questions from the section above. Three: if the LLM is misplaced, design the replacement in four layers (official source, derived catalogue, synonyms table, review queue). Four: run the replacement in parallel with the current system for two weeks, comparing results case by case. Five: once the deterministic function beats the LLM on accuracy and cost, switch the LLM off. The whole process rarely takes more than three weeks if you have access to the code and to the domain expert.
Does determinism also apply to autonomous agents?
Yes, and that's where it matters most. An autonomous agent that takes actions (sends emails, moves money, updates records) with no human in the loop multiplies the cost of every bad decision. The right architecture uses the agent to reason over variable input and deterministic code to execute any action with side effects. The agent proposes, the deterministic function disposes. If the agent decides to assign a TARIC code and clear the shipment without going through the catalogue layer, you don't have an autonomous agent. You have an agent out of control.
Closing
The expensive mistake isn't choosing the wrong model. It's putting the model in the wrong place.
Every month your workflow sends €200 to OpenAI to classify codes published in a free European Commission XML file is a month you pay for a worse result. Every quarter the senior broker spends correcting the LLM's answers by hand instead of maintaining a synonyms table is a quarter's salary spent on a problem that shouldn't exist. Every audit where you can't reproduce a classification from six months ago, because that version of the model is no longer available, is a hidden liability in your architecture.
Determinism first. The model goes where determinism can't reach, and not before.
If you have a workflow with an LLM you suspect is misplaced and want to check which steps determinism wins, book a 30-minute session to run the six-criteria diagnosis together and leave with a concrete list of which steps can move to code and which justify keeping the model. If you'd rather first explore which model fits each type of task, the AI models by task page has the full map.
Related reading
- Semantic Scale: when to move up from dictionary to embedding to LLM
- The real cost of an LLM architecture: the bill nobody shows you
- Auditing legacy AI workflows against six criteria
Sources
- European Commission, official TARIC database: ec.europa.eu/taxation_customs/dds2/taric
- INE, CNAE 2025 classification: ine.es/dyngs/INEbase/es/operacion.htm?c=Estadistica_C&cid=1254736177032
- OpenAI, API pricing in force in March 2026: openai.com/api/pricing
- Anthropic, API pricing in force in March 2026: anthropic.com/pricing
