Home/Insights/Where AI Still Needs Deterministic Rules

Where AI Still Needs Deterministic Rules: At Work, In Software, As a Customer

Quick Answer

AI models are built to vary. That helps when you want a draft and hurts when you want a decision. Wherever the same input must always give the same output - prices, eligibility, approvals, permissions, policy answers - the logic belongs in rules and code, and the model stays in the parts where variation is harmless.

The evidence is specific. At temperature 0, 1,000 identical requests returned 80 different answers in one test. In a separate study, accuracy varied by up to 15% between runs of models set to be deterministic. Insurance underwriters priced identical cases 55% apart on average, more than five times what their executives expected, so humans are not the consistent baseline either. And a step that is 95% reliable repeated ten times succeeds about 60% of the time. What follows covers each arena and the seven assumptions that push teams the wrong way, including how an interface should behave.

Author: May Mor - Efficiency Leader. I help operators align their people, systems, and processes so growth scales the business instead of breaking it. M.Sc in AI, 10+ years across fintech, digital banking and adtech.

What Deterministic Means, and Where Variation Is Fine

Deterministic means the same input produces the same output, every time. A calculator is deterministic. A language model is not, by design: it picks each next word from a probability distribution, which is what lets it write, summarize and improvise. Remove that and you remove most of what makes it useful.

So the question for a company is not whether AI is reliable. It is which parts of the work may vary and which may not. The table is the short version.

The workVariationWhere it belongs
Drafting an email, a first summary, a list of ideasFine, often usefulThe model
Reading a document, classifying a request, extracting fieldsFine if a rule or a person checks the resultThe model, then a check
Price, discount limit, eligibility, refund amountNot fine: same case, same answerRules or code
Approvals, permissions, releasing a paymentNot fineRules or code, with a log
Policy wording a customer relies onNot fine: it is a commitmentApproved text, looked up and not regenerated

At Work: People Are Not the Consistent Baseline

Assumption one: human judgment is the steady reference point, and AI is the wobbly newcomer. The research says companies already have a consistency problem before any model arrives.

Daniel Kahneman and colleagues ran a noise audit at an insurer: 48 underwriters priced the same realistic cases. Executives expected a difference of about 10% between two underwriters. The typical difference was 55% of the average premium, more than five times as large. They named the chance variability of judgments noise, and described it as an invisible tax on the bottom line.

A formula does not get tired, anchored by the last case, or lenient before lunch. A meta-analysis by Grove and colleagues of 136 studies of human health and behavior found mechanical prediction about 10% more accurate on average, with clinical judgment substantially more accurate in only a small minority of studies (6% to 16%, depending on the analysis). Those studies cover prediction tasks in health and behavior, so read them as strong evidence for rules in repeatable judgments, not for every kind of work.

This is the old case for scorecards, checklists and approval limits, and AI does not retire it. A model used as the decision maker adds its own variation on top of the human variation already there. My recommendation is to write the rule first. If a decision can be described as a rule - a limit, a threshold, a list of required fields - let the model prepare the inputs by reading the document and extracting the fields, and let the rule decide. If you cannot write the rule, you also cannot audit what the model decided.

In Software: Three Assumptions That Break

Assumption two: temperature 0 makes the model deterministic. Temperature controls how much randomness the model uses when picking words, and 0 means always take the most likely one. It should give identical output. It does not. Thinking Machines Lab sent the same request 1,000 times at temperature 0 and received 80 unique completions. They traced it to batch size: the same request gives slightly different internal numbers depending on how many other requests the server is handling at that moment, and server load changes constantly. With batch-invariant kernels, all 1,000 outputs were identical. The variation comes from how the service is run, not only from the model.

A separate study by Atil and colleagues ran five LLMs configured to be deterministic across eight tasks, ten runs each. Accuracy varied by up to 15% between runs, the gap between best and worst possible performance reached up to 70%, and no model gave repeatable accuracy on every task.

Assumption three: a high multiple-choice score means the model is reliable. Most public benchmarks are multiple-choice tests. Zheng and colleagues tested 20 LLMs on three benchmarks and found the models change their answer when the options are reordered, because they favor certain option labels, such as A. A model that gives a different answer to the same question with the same options in a different order is giving you a measurement of the format along with a measurement of the knowledge. A simple check for your own use case: shuffle the options and run it again.

Assumption four: if each step is 95% reliable, the whole task is 95% reliable. It is not, because errors multiply. The table is plain arithmetic and assumes the steps are independent and that one failed step fails the task.

Steps in the taskEach step 99% reliableEach step 95% reliable
199%95%
595%77%
1090%60%
2082%36%

The measured version is τ-bench, a 2024 benchmark from Sierra Research that simulates customer-service tasks with tools and written policies. It introduced a metric called pass^k: the probability of succeeding on all k repeated runs of the same task. The best function-calling agents it tested succeeded on fewer than half of the tasks, and on the retail tasks the chance of succeeding on all eight repeated runs was below 25%. The headline success rate and the repeat-run rate are two different numbers, and a process that serves customers needs the second one.

Anthropic draws the design line in its guidance on building agents. A workflow orchestrates models and tools through predefined code paths. An agent lets the model direct its own process. Their advice is to use the simplest solution that works, and to add agentic freedom only when the task needs it, because it trades latency and cost for performance. For a task with known, repeatable steps and little tolerance for variance, a fixed workflow is the sound default.

In practice, I would hold software that uses a model to five habits: test with repeated runs and report the share that pass every time, not the average; pin the model version and log every input and output; validate model output against a schema before anything uses it; keep every action with a consequence (payments, deletions, emails, permission changes) behind ordinary code with its own checks; and write the rules the model must not override as code, not as a sentence in the prompt.

As a Customer: The Same Question Needs the Same Answer

A customer asks a question once and relies on the answer. In February 2024 the Civil Resolution Tribunal of British Columbia decided Moffatt v. Air Canada. The airline's website chatbot told a customer he could claim a bereavement fare retroactively within 90 days. The airline's actual policy said it would not refund bereavement travel after booking. The tribunal found that Air Canada "did not take reasonable care to ensure its chatbot was accurate" and ordered it to pay CAD 812 in damages, interest and fees. The amount was small. The decision holds the company responsible for what its own chatbot told a customer.

Customers are wary for related reasons. A Gartner survey of 5,728 customers in December 2023 found that 64% would prefer companies did not use AI in customer service, and 53% would consider switching to a competitor if they learned a company planned to. The top concerns were that reaching a person would get harder, that AI would displace jobs, and that AI would give wrong answers.

Assumption five: a chatbot's answer is a suggestion. To the customer it is the company's position. That is the reason policy answers, prices and eligibility should come from the approved source and not from a fresh generation each time. A customer who asks twice and gets two answers has no way of knowing which one counts.

Assumption six: more choices serve the customer, or the reverse. This is the multiple-choice fallacy on the customer side, and both camps are overconfident. The famous jam study by Iyengar and Lepper reported that 3% of shoppers who saw 24 jams used a discount voucher, against 30% of those who saw six. A meta-analysis by Scheibehenne, Greifeneder and Todd later pooled 63 conditions from 50 experiments (5,036 participants) and found the average effect of too many options was virtually zero, with large variation between studies. Whether fewer options help depends on the situation, so the number of options is something to test in your own product.

It is also something to fix. If an AI system generates the list of options, the plans or the price shown to a customer fresh on every visit, the customer who reloads sees a different set and cannot tell which one is real. Decide the choice set with rules, test it, and keep it stable for the same customer in the same situation.

In the Interface: Users Learn Your Product, So It Cannot Reshuffle

Assumption seven: because AI can generate an interface for each person, it should. People learn an interface the way they learn a building: where things are, what a button does, what happens next. Consistency and standards is the fourth of Jakob Nielsen's ten usability heuristics, and Jakob's Law gives the reason: users spend most of their time in other products and expect yours to work the same way. An interface assembled fresh for each session breaks both, because nothing the user learned yesterday is guaranteed to hold today.

The research on interfaces that change by themselves is older than generative AI, and it points one way. Findlater and McGrenere (2004) compared three versions of a menu: static, adaptable (the user moves items), and adaptive (the system reorders items based on use). The static menu was significantly faster than the adaptive one, and most participants preferred the adaptable menu, where the user controls the change. A system that rearranges things on its own cost speed, and given the choice, users picked control.

Microsoft Research's guidelines for human-AI interaction (Amershi et al., CHI 2019: 18 guidelines, tested with 49 practitioners across 20 AI products) say the same thing twice. Guideline 14 is to update and adapt cautiously: limit disruptive changes, and consider the scale and rate of change. Guideline 18 is to notify users about changes, so they can recalibrate their expectations.

In design terms, these are my recommendations rather than findings from those studies. Keep layout, navigation and control placement fixed, and let the model fill defined content slots inside them. Generate button labels and error messages once, review them, and ship them as ordinary text instead of generating them at runtime. When a model's answer can vary, present it as a draft the user can accept, edit or regenerate on purpose, and store what they accepted so it does not change the next time they open it. Say what changed after any model update that alters behavior. Test the interface for consistency the way you test the logic: the same input should render the same screen.

Where to Draw the Line

The pattern that holds across work, software and customers is the same: the model reads and proposes, the rules decide, and the code acts.

  1. List the decisions. Write down every place a model output could change a price, a promise, a permission or a payment.
  2. Write the rule for each one. If it can be a threshold, a table or a checklist, make it one, and make it testable.
  3. Give the model the reading and drafting. Extracting fields, classifying requests, summarizing and drafting are where variation is cheap.
  4. Put a deterministic check between the model and the action. Schema validation, limits, an approval step, a log.
  5. Measure consistency, not only accuracy. Run the same case many times and track how often the answer is the same.

None of this is an argument against using AI. It is an argument for being specific about where its variation costs nothing and where it costs money, a promise or a customer.

Sources

May Mor
About the author

May Mor

Efficiency Leader. I help operators align their people, systems, and processes so growth scales the business instead of breaking it. M.Sc in AI, 10+ years across fintech, digital banking and adtech. Full bio →

If you are deciding where AI belongs in a process, and where rules must stay:
Scale Readiness Assessment Scoped to your budget and timeline - the diagnosis before any engagement
Book a 30-min intro call 30 min to know if this fits your situation