Semantic Scale: seven levels to assign each task to the right mechanism
Every tool is an answer to a question. Using the wrong answer costs real money.
The cost per million tokens between a frontier model and a lightweight model varies by a factor of 30x to 100x depending on the provider and the context window. The same business workflow can cost €40,000 a month or €400, depending on which task is assigned to which mechanism. It isn't a question of provider, of negotiating rates or of waiting for Anthropic or OpenAI to cut prices. It's a question of design.
Most of the AI architectures I see in service companies share the same defect: they use the most expensive model for every task in the workflow, including the ones twenty lines of Python would solve. You pay for the multi-step reasoning of a frontier model to check that an IBAN starts with ES and has 24 characters. You pay to read an entire PDF to pull out a date that was already in the file name. You pay to generate a classification that already exists in a static table of 300 rows.
The Semantic Scale is a method for assigning each task in the workflow to the level of mechanism it needs, and not one level more. Seven levels, from pure deterministic code to pure human decision. Each level solves a different type of problem, and the cost per query differs between them by several orders of magnitude.
This isn't theory. It's the only way to cut the bill without cutting functionality.
The seven levels, one by one
The levels aren't ordered by technical sophistication. They're ordered by how much indeterminacy the input is allowed to have. The more rigid the input and the more predictable the result, the lower the level. The more ambiguous the input or the harder the criterion is to encode, the higher.
Level 1: pure deterministic code
Regex, format validation, arithmetic, string comparison, hashing. Structured input, predictable output, negligible computational cost. Examples: validating the check digit of a Spanish tax ID (NIF), checking that an IBAN complies with the ISO 13616 standard, calculating 21% VAT on a taxable base, adding up the lines of an invoice.
Cost per query: practically zero. Latency: microseconds. Auditability: total. Risk of error: none, if the code is well written and tested.
Rule: if a task can be described with a formal pattern or a closed formula, it doesn't go above level 1. Ever.
Level 2: rules and lookup tables
Dictionaries, taxonomies, encoded decision trees, translation maps. Still deterministic, but with domain knowledge packaged as data. Examples: mapping a CNAE code (Spain's official classification of economic activities) to its sector, a postcode to its province, an IBAN to its issuing bank, an LEI code to the company's legal name.
The logic is the same as level 1. The difference is that the answer isn't calculated but looked up in a table built beforehand. Cost per query: microseconds if the table is in memory, milliseconds if it's on disk or in a database.
Example: a gestoría (a Spanish firm that handles tax filings and bookkeeping for small businesses) processes 4,000 invoices a month. 78% of the suppliers are already in its database. Matching the invoice's tax ID against that table resolves expense category, ledger account and cost centre without touching a single LLM token. Only the remaining 22% moves up to the next level.
Level 3: classical machine learning models
Trained classifiers (logistic regression, random forest, gradient boosting), embeddings plus vector search, standard OCR, similarity-based duplicate detection. They need prior training and labelled data, but inference is cheap and fast.
Examples: classifying an incoming email into 12 categories with a model trained on 6,000 historical emails, extracting text from a scanned PDF with open-source OCR, finding the supplier closest to "Suminstros Herrera SL" (typo included) in a database of 12,000 suppliers using embeddings.
Cost per query: between 0.001 and 0.01 US cents on self-hosted models. Latency: tens of milliseconds. Quality depends on the training dataset and degrades if the input distribution changes.
Critical note: level 3 has a bad reputation in 2026 because it looks "less modern" than an LLM. That's a mistake. A classifier trained on your own data is usually more accurate, cheaper and faster than a lightweight LLM for the same closed classification task.
Level 4: lightweight LLM
Small models optimised for low-complexity, high-speed tasks. In 2026 the category includes Claude Haiku, GPT-4o mini, Gemini Flash, Mistral Small and Llama 3.1 8B. Typical input cost: $0.15 to $0.50 per million tokens. Typical output cost: $0.60 to $2.50 per million tokens.
Used well, they handle structured extraction from low-variability documents, prompt-based classification when training your own model doesn't pay off, paragraph summaries, email rewrites, and generating JSON fields from free text.
Example: extracting the ten usual fields of an invoice (date, number, taxable base, VAT, total, issuer tax ID, recipient tax ID, description, payment method, due date) from the OCR already done at level 3. The lightweight LLM returns JSON validated against a schema. Real cost per invoice processed: between 0.04 and 0.10 US cents.
Level 5: frontier LLM
Large models capable of multi-step reasoning, long context, non-trivial code and following complex instructions. Claude Opus, GPT-5, Gemini Ultra. Typical input cost: $3 to $15 per million tokens. Typical output cost: $15 to $75 per million tokens.
They're justified when the task demands:
- Non-linear reasoning. Spotting an anomaly in a series of accounting entries that you can't see by looking at any single entry.
- Complex writing with judgement. A technical tender specification, an audit report, a draft legal opinion.
- Code of more than 50 lines with dependencies between functions.
- Long-context comprehension above 50,000 tokens, where the lightweight models get lost.
If the task doesn't fit any of those cases, moving up to level 5 is throwing money away.
Level 6: frontier LLM with mandatory human review
Same engine as level 5, but the output isn't executed or sent to the client until a person signs it off. The review isn't decoration: it's a control of legal, contractual or clinical liability.
Examples: a contract draft generated by an LLM that a lawyer reviews and edits before sending, a medical report drafted by an LLM that the physician approves and signs, a reply to a consumer complaint that the customer service manager validates before it goes out under the company logo.
Cost per query: the level 5 cost plus the human review time (between 3 and 45 minutes depending on complexity, at the professional's hourly cost).
Level 7: pure human decision
No AI on the critical path. The person decides with their own judgement, their experience and whatever information they consider relevant. AI may appear beforehand as an information assistant, but the decision isn't recorded as generated by a model, isn't automated and isn't delegated.
Examples: dismissing an employee, accepting or rejecting a client, setting the year's pricing strategy, deciding whether to litigate or settle a lawsuit, choosing a business partner, writing the email to a client with whom there's a real emotional conflict.
Level 7 doesn't disappear with AI. It narrows, but it doesn't disappear. We'll come back to this.
A real document workflow mapped to the seven levels
The best way to understand the Semantic Scale is to apply it to a specific workflow. Take the processing of incoming invoices at an accounting firm that manages 90 clients and receives 4,000 invoices a month.
| Step | Task | Level | Specific mechanism |
|---|---|---|---|
| 1 | Check the file is PDF, JPG or EML and no larger than 15 MB | 1 | Validation by extension and size |
| 2 | Extract the issuer's tax ID if it appears in plain text or in the file name | 1 | Regex [A-Z]?\d{7,8}[A-Z]? plus check digit |
| 3 | Map the tax ID to a known supplier and its ledger account | 2 | Lookup in the client's table of 12,400 suppliers |
| 4 | If scanned, run OCR on the images | 3 | OCR with an open model tuned to Spanish invoices |
| 5 | Extract the ten structured invoice fields | 4 | Lightweight LLM with extraction prompt and JSON schema validation |
| 6 | Categorise the expense according to the client's chart of accounts | 3 or 4 | Classifier trained on history, or lightweight LLM if the client is new |
| 7 | Check arithmetic consistency (base plus VAT equals total) | 1 | Direct arithmetic with a €0.02 tolerance |
| 8 | Flag a suspicious invoice (amount out of range, unusual description, VAT not applicable to the client's tax regime) | 5 | Frontier LLM with the context of the last 24 invoices from the same supplier |
| 9 | If the amount exceeds €3,000 or the anomaly is serious, escalate to review | 6 | Accountant reviews and approves in the internal interface |
| 10 | Decide whether to open a formal case for possible fraud or supplier error | 7 | Head of the firm decides case by case |
Of the ten steps, two use a lightweight LLM, one uses a frontier LLM, one requires human review and only the last is a pure human decision. The other six steps are solved with code, tables and classical models.
In plain terms: 60% of the workflow never sees a single LLM token. That 60% is where the savings live.
Cost per query: level 4 against level 5 on an identical case
Let's get to the number. Take step 5 of the workflow above: extracting the ten structured fields of an invoice that has already been digitised. The input is the invoice's plain text after OCR, about 2,000 tokens. The output is JSON of about 300 tokens.
Scenario A: lightweight LLM (level 4).
At a typical rate of $0.25 per million input tokens and $1.25 per million output tokens:
- Input: 2,000 × 0.25 / 1,000,000 = $0.0005
- Output: 300 × 1.25 / 1,000,000 = $0.000375
- Cost per invoice: $0.000875. Rounded, 0.09 US cents.
For 4,000 invoices a month: $3.50. For the 48,000 invoices a year: $42.
Scenario B: frontier LLM (level 5).
At a typical rate of $15 per million input tokens and $75 per million output tokens:
- Input: 2,000 × 15 / 1,000,000 = $0.03
- Output: 300 × 75 / 1,000,000 = $0.0225
- Cost per invoice: $0.0525. Rounded, 5.25 US cents.
For 4,000 invoices a month: $210. For the 48,000 invoices a year: $2,520.
Difference: 60x. The frontier model costs 60 times more to do exactly the same job.
Now the important question: do you gain quality? In structured extraction from low-variability documents, with a JSON schema validated by level 1 code (checking that the amount is a number, that the date is valid, that base plus VAT matches the total), the accuracy gap between a modern lightweight LLM and a frontier model is in the order of 1% to 3%. At an accounting firm handling 4,000 invoices a month, that 1-3% means between 40 and 120 invoices a month that the lightweight model gets wrong.
Fix: the lightweight model does the extraction, level 1 code validates the result, and only if validation fails is the task retried with the frontier model. Instead of paying $210 a month, you pay between $4 and $15. Final quality is equivalent because the lightweight model's failures are caught and reprocessed.
This structure, "cheap first, expensive only when the cheap one fails", is the central pattern of the Semantic Scale. It's called ascending fallback, and it's where most of the real savings live.
When liability forces you up to level 6
Moving down a level saves money. Moving up a level covers liability. Both moves are legitimate, but they're justified by different criteria.
An automated workflow goes to level 6 (mandatory human review) when at least one of these three factors is present:
- Direct legal consequence of the output. Contracts, legal opinions, official communications to public authorities, reports that go out signed. If the text creates an obligation or extinguishes a right, it can't go out without a human signature.
- Irreversible financial consequence above a threshold. Payments, transfers, credit approvals, spending authorisations. Internal policy sets the threshold, but below €500 many companies accept automation; above €3,000 almost none do any more.
- Clinical or safety consequence. Medical diagnosis, dosage, structural safety report, technical opinion that determines an intervention on a patient or a building. Full automation is ruled out by the regulatory framework, not by business judgement.
Outside those three, moving up to level 6 is usually excess caution that eats staff hours with no real gain. A typical case: manually reviewing the email summaries the AI generates for the CRM. If a summary is wrong, it gets corrected when someone reads the original email. The error has no consequence. The systematic review costs the salesperson an hour a day. Net gain: negative.
Rule: level 6 is justified by the size and irreversibility of the consequence of an error, not by the nerves of the systems architect.
There's also a reverse route. Many workflows start at level 6 and drop to level 5 once the organisation has built up enough history to trust them. A law firm that reviews 100% of LLM-generated drafts during the first six months can move to reviewing only a 10% sample, once it has evidence that the error rate is below 2% for a certain type of document. The review doesn't disappear. It gets sampled.
Why level 7 doesn't disappear
There's a recurring fantasy in AI architecture presentations: the end-to-end workflow with no human at any node. The document comes in, the system decides, the response goes out, the client receives it. Zero friction. Zero human cost.
No.
Level 7 doesn't disappear, for four structural reasons.
First: there are decisions an organisation can't delegate without destroying its own legitimacy. Hiring and firing. Accepting and rejecting clients. Setting strategic prices. Changing partners. Litigating or settling. These decisions have a political, ethical and relational component. If they're automated, the organisation loses authority inside and out. Nobody signs a €200,000-a-year contract with a supplier whose decision to take them on was made by an LLM with no human behind it.
Second: there are decisions that are irreversible and infrequent. Automating them isn't worth it. The cost of designing, testing, maintaining and auditing a system to make a decision that comes up twice a year far exceeds the cost of two people sitting down for an afternoon and deciding.
Third: there are tasks where the value lies in the human process, not in the result. An exit interview with an employee who's leaving. An apology call to an angry client. A renewal negotiation with a long-standing partner. Automate it and the formal result appears, but the relational value that justified the task disappears. The task stops doing its job.
Fourth: there are risks no insurer covers if the process is automatic. Critical diagnosis, the decision to administer a certain drug, approval of a surgical procedure, a structural opinion on an occupied building. The insurer requires an identifiable human signature. Without it, there's no policy.
Practical consequence: level 7 narrows, but it stays. A well-designed workflow doesn't eliminate pure human decisions. It concentrates them in fewer people, with more context and better information prepared beforehand by levels 1 to 6. That's what lowers the cost without destroying accountability.
Checklist for designing the Semantic Scale in your workflow
This is the operating procedure, condensed. It applies to any business workflow that runs expensive AI today and wants to lower the cost without losing quality.
- List every atomic task in the workflow. An atomic task is the smallest unit that produces a verifiable output. Not "process invoice", but "extract tax ID", "validate tax ID", "map tax ID to supplier", "extract amount", "validate arithmetic". If a task on the list has more than one output, it isn't atomic yet.
- For each task, write the success criterion in one sentence. If you can't write it in one sentence, the task still isn't properly scoped.
- Ask: is there a formal pattern, a formula or a closed rule that solves the task? If yes, the task is level 1. Moving up is forbidden.
- Ask: is there a table, a dictionary or a knowledge map that solves the task with a lookup? If yes, and the table can be built and maintained, the task is level 2.
- Ask: do I have enough labelled historical data to train my own classifier? If yes, and the task is classification or similarity, the task is level 3.
- Ask: is the task structured extraction, prompt-based classification, a short summary or a simple rewrite? If yes, and the volume justifies the cost, it's level 4. With fallback to level 5 only when level 4 fails validation.
- Ask: does the task demand multi-step reasoning, long context, complex writing or non-trivial code? If yes, level 5.
- Ask: does an error have a legal, irreversible financial or clinical consequence? If yes, add mandatory human review at the end of level 5. Result: level 6.
- Ask: does the decision set organisational policy, involve a relationship with a person or carry insurance risk? If yes, the task isn't automated. Level 7.
- Measure the real cost per query and per month of every task in the workflow. Don't estimate, measure. With token logs if it's an LLM, with execution time if it's your own code, with human time if there's review. Without a metric, the scale degrades because nobody knows whether costs are going up or down.
Applying this checklist to an existing workflow takes between two and five working days, depending on the number of atomic tasks and access to cost data. The typical result in accounting firms, law firms and administrative service companies is a 55% to 80% reduction in the AI bill with no measurable loss of final quality.
Common mistakes when implementing the Semantic Scale
The theory is simple. The execution has known traps.
- Mistake 1: defaulting to the LLM because "it's more modern". Symptom: the first version of the workflow uses a frontier LLM for everything, format validation included. Fix: apply the checklist above with discipline. No task moves up a level if the level below solves it.
- Mistake 2: using a frontier LLM for closed classification. Symptom: the task is assigning a label from a finite set (for example, an expense category out of 42 options) and it's solved by asking the LLM to pick one. Fix: if there's labelled historical data, train a level 3 classifier. If there isn't yet, use a lightweight LLM while it builds up, and drop to level 3 once you have 3,000 labelled examples.
- Mistake 3: long prompts with context repeated on every call. Symptom: every LLM call includes 8,000 tokens of instructions and examples that never change. Fix: prompt caching if the provider supports it (it cuts the cost of the cached part to 10%-25%), or moving to a model with enforced JSON schema that halves the instructions.
- Mistake 4: human review of everything, just in case. Symptom: the workflow goes through human review at every step, tax ID validation included. Fatigue guaranteed. Fix: apply the level 6 criterion only where there's a legal, irreversible financial or clinical consequence. Everywhere else, random sampling of 5% to 10% for quality control, without blocking the workflow.
- Mistake 5: not measuring the cost per query. Symptom: the LLM provider's invoice arrives at the end of the month and it's a surprise. Nobody knows which task consumed what. Fix: instrument the workflow from day one with a token counter per task, store it in a database, review it weekly. Without a metric, the scale degrades.
- Mistake 6: ignoring perceived latency. Symptom: the workflow is cheap but takes 40 seconds per document because it chains six LLM calls in series. Fix: parallelise the independent calls, use lightweight models for the high-volume tasks, and keep the frontier model only for what demands reasoning. Latency drops from 40 seconds to 6 without touching the cost.
Frequently asked questions
How long does it take to redesign a workflow with the Semantic Scale?
Between two and five working days for a workflow of 8 to 15 atomic tasks, if there's access to the cost data and the logs. The result shows up in the first LLM provider invoice after deployment.
Does the Semantic Scale also work for content generation workflows, not just processing?
Yes. The logic is identical. An automated blog, for example, can handle source research with vector search (level 3), the draft with a lightweight LLM (level 4), the style review with a frontier LLM (level 5) and final approval with a human (level 6). There's no need to use the most expensive model at every step.
Which specific models fall into the lightweight LLM category in 2026?
Claude Haiku, GPT-4o mini, Gemini Flash, Mistral Small, Llama 3.1 8B and self-hosted equivalents. The category is defined by price (below $1 per million input tokens) and by task profile (extraction, prompt-based classification, short summaries), not by provider.
How do you justify to management the investment of redesigning a workflow that already works?
With the number. If the current workflow spends €8,000 a month on LLM tokens and the redesign brings that down to €1,400 a month with no measurable loss of quality, the redesign investment (between €6,000 and €15,000 according to the usual market ranges published by Vertebra Gestión and ASD Solutions) pays for itself in under three months.
Can the Semantic Scale be applied without an in-house technical team?
It can be designed without an in-house team, but it has to be implemented by someone who can code. Designing the scale means analysing the workflow and deciding the level of each task. Implementing it means writing the level 1 and 2 code, training the level 3 models if there are any, integrating the LLM calls and setting up the human controls. The first part is done in a hands-on workshop; the second takes days or weeks of development.
What happens if an LLM provider raises prices or retires a model?
If the scale is well designed, switching lightweight LLM provider affects only 30%-40% of the workflow. The rest (deterministic code, tables, classical models, human review) is independent of the provider. That's one of the reasons the Semantic Scale reduces risk as well as cost.
Closing
The AI bill in service companies doesn't grow because of the providers. It grows because the design uses the same expensive mechanism for every task in the workflow. Changing that doesn't require new technology or better negotiated rates. It requires looking at each task in the workflow and deciding which level it needs, and not one level more.
The seven levels aren't a theoretical model. They're an operating rule that takes two working days per workflow to apply and cuts the bill by 55% to 80% on the first invoice afterwards. Final quality doesn't drop because failures at the cheap levels are caught by code and reprocessed only when needed. Accountability doesn't degrade because the steps with legal, financial or clinical consequences stay at level 6 or 7. And pure human decisions keep happening where they should, with less noise around them because the six levels below have already filtered out whatever didn't need human judgement.
If your current AI workflow sits in the €3,000 to €40,000 range of monthly token spend and keeps growing without clear control, the Semantic Scale is the shortest route to bringing it down without losing functionality. You can go deeper into the method at /escala-semantica or book a 30-minute session to review your current workflow, list its atomic tasks and estimate the specific savings that would apply in your case. Free of charge.
Related reading
