What Deterministic Means, and Where Variation Is Fine
Deterministic means the same input produces the same output, every time. A calculator is deterministic. A language model is not, by design: it picks each next word from a probability distribution, which is what lets it write, summarize and improvise. Remove that and you remove most of what makes it useful.
So the question for a company is not whether AI is reliable. It is which parts of the work may vary and which may not. The table is the short version.
| The work | Variation | Where it belongs |
|---|---|---|
| Drafting an email, a first summary, a list of ideas | Fine, often useful | The model |
| Reading a document, classifying a request, extracting fields | Fine if a rule or a person checks the result | The model, then a check |
| Price, discount limit, eligibility, refund amount | Not fine: same case, same answer | Rules or code |
| Approvals, permissions, releasing a payment | Not fine | Rules or code, with a log |
| Policy wording a customer relies on | Not fine: it is a commitment | Approved text, looked up and not regenerated |
At Work: People Are Not the Consistent Baseline
Assumption one: human judgment is the steady reference point, and AI is the wobbly newcomer. The research says companies already have a consistency problem before any model arrives.
Daniel Kahneman and colleagues ran a noise audit at an insurer: 48 underwriters priced the same realistic cases. Executives expected a difference of about 10% between two underwriters. The typical difference was 55% of the average premium, more than five times as large. They named the chance variability of judgments noise, and described it as an invisible tax on the bottom line.
A formula does not get tired, anchored by the last case, or lenient before lunch. A meta-analysis by Grove and colleagues of 136 studies of human health and behavior found mechanical prediction about 10% more accurate on average, with clinical judgment substantially more accurate in only a small minority of studies (6% to 16%, depending on the analysis). Those studies cover prediction tasks in health and behavior, so read them as strong evidence for rules in repeatable judgments, not for every kind of work.
This is the old case for scorecards, checklists and approval limits, and AI does not retire it. A model used as the decision maker adds its own variation on top of the human variation already there. My recommendation is to write the rule first. If a decision can be described as a rule - a limit, a threshold, a list of required fields - let the model prepare the inputs by reading the document and extracting the fields, and let the rule decide. If you cannot write the rule, you also cannot audit what the model decided.
In Software: Three Assumptions That Break
Assumption two: temperature 0 makes the model deterministic. Temperature controls how much randomness the model uses when picking words, and 0 means always take the most likely one. It should give identical output. It does not. Thinking Machines Lab sent the same request 1,000 times at temperature 0 and received 80 unique completions. They traced it to batch size: the same request gives slightly different internal numbers depending on how many other requests the server is handling at that moment, and server load changes constantly. With batch-invariant kernels, all 1,000 outputs were identical. The variation comes from how the service is run, not only from the model.
A separate study by Atil and colleagues ran five LLMs configured to be deterministic across eight tasks, ten runs each. Accuracy varied by up to 15% between runs, the gap between best and worst possible performance reached up to 70%, and no model gave repeatable accuracy on every task.
Assumption three: a high multiple-choice score means the model is reliable. Most public benchmarks are multiple-choice tests. Zheng and colleagues tested 20 LLMs on three benchmarks and found the models change their answer when the options are reordered, because they favor certain option labels, such as A. A model that gives a different answer to the same question with the same options in a different order is giving you a measurement of the format along with a measurement of the knowledge. A simple check for your own use case: shuffle the options and run it again.
Assumption four: if each step is 95% reliable, the whole task is 95% reliable. It is not, because errors multiply. The table is plain arithmetic and assumes the steps are independent and that one failed step fails the task.
| Steps in the task | Each step 99% reliable | Each step 95% reliable |
|---|---|---|
| 1 | 99% | 95% |
| 5 | 95% | 77% |
| 10 | 90% | 60% |
| 20 | 82% | 36% |
The measured version is τ-bench, a 2024 benchmark from Sierra Research that simulates customer-service tasks with tools and written policies. It introduced a metric called pass^k: the probability of succeeding on all k repeated runs of the same task. The best function-calling agents it tested succeeded on fewer than half of the tasks, and on the retail tasks the chance of succeeding on all eight repeated runs was below 25%. The headline success rate and the repeat-run rate are two different numbers, and a process that serves customers needs the second one.
Anthropic draws the design line in its guidance on building agents. A workflow orchestrates models and tools through predefined code paths. An agent lets the model direct its own process. Their advice is to use the simplest solution that works, and to add agentic freedom only when the task needs it, because it trades latency and cost for performance. For a task with known, repeatable steps and little tolerance for variance, a fixed workflow is the sound default.
In practice, I would hold software that uses a model to five habits: test with repeated runs and report the share that pass every time, not the average; pin the model version and log every input and output; validate model output against a schema before anything uses it; keep every action with a consequence (payments, deletions, emails, permission changes) behind ordinary code with its own checks; and write the rules the model must not override as code, not as a sentence in the prompt.
As a Customer: The Same Question Needs the Same Answer
A customer asks a question once and relies on the answer. In February 2024 the Civil Resolution Tribunal of British Columbia decided Moffatt v. Air Canada. The airline's website chatbot told a customer he could claim a bereavement fare retroactively within 90 days. The airline's actual policy said it would not refund bereavement travel after booking. The tribunal found that Air Canada "did not take reasonable care to ensure its chatbot was accurate" and ordered it to pay CAD 812 in damages, interest and fees. The amount was small. The decision holds the company responsible for what its own chatbot told a customer.
Customers are wary for related reasons. A Gartner survey of 5,728 customers in December 2023 found that 64% would prefer companies did not use AI in customer service, and 53% would consider switching to a competitor if they learned a company planned to. The top concerns were that reaching a person would get harder, that AI would displace jobs, and that AI would give wrong answers.
Assumption five: a chatbot's answer is a suggestion. To the customer it is the company's position. That is the reason policy answers, prices and eligibility should come from the approved source and not from a fresh generation each time. A customer who asks twice and gets two answers has no way of knowing which one counts.
Assumption six: more choices serve the customer, or the reverse. This is the multiple-choice fallacy on the customer side, and both camps are overconfident. The famous jam study by Iyengar and Lepper reported that 3% of shoppers who saw 24 jams used a discount voucher, against 30% of those who saw six. A meta-analysis by Scheibehenne, Greifeneder and Todd later pooled 63 conditions from 50 experiments (5,036 participants) and found the average effect of too many options was virtually zero, with large variation between studies. Whether fewer options help depends on the situation, so the number of options is something to test in your own product.
It is also something to fix. If an AI system generates the list of options, the plans or the price shown to a customer fresh on every visit, the customer who reloads sees a different set and cannot tell which one is real. Decide the choice set with rules, test it, and keep it stable for the same customer in the same situation.
In the Interface: Users Learn Your Product, So It Cannot Reshuffle
Assumption seven: because AI can generate an interface for each person, it should. People learn an interface the way they learn a building: where things are, what a button does, what happens next. Consistency and standards is the fourth of Jakob Nielsen's ten usability heuristics, and Jakob's Law gives the reason: users spend most of their time in other products and expect yours to work the same way. An interface assembled fresh for each session breaks both, because nothing the user learned yesterday is guaranteed to hold today.
The research on interfaces that change by themselves is older than generative AI, and it points one way. Findlater and McGrenere (2004) compared three versions of a menu: static, adaptable (the user moves items), and adaptive (the system reorders items based on use). The static menu was significantly faster than the adaptive one, and most participants preferred the adaptable menu, where the user controls the change. A system that rearranges things on its own cost speed, and given the choice, users picked control.
Microsoft Research's guidelines for human-AI interaction (Amershi et al., CHI 2019: 18 guidelines, tested with 49 practitioners across 20 AI products) say the same thing twice. Guideline 14 is to update and adapt cautiously: limit disruptive changes, and consider the scale and rate of change. Guideline 18 is to notify users about changes, so they can recalibrate their expectations.
In design terms, these are my recommendations rather than findings from those studies. Keep layout, navigation and control placement fixed, and let the model fill defined content slots inside them. Generate button labels and error messages once, review them, and ship them as ordinary text instead of generating them at runtime. When a model's answer can vary, present it as a draft the user can accept, edit or regenerate on purpose, and store what they accepted so it does not change the next time they open it. Say what changed after any model update that alters behavior. Test the interface for consistency the way you test the logic: the same input should render the same screen.
Where to Draw the Line
The pattern that holds across work, software and customers is the same: the model reads and proposes, the rules decide, and the code acts.
- List the decisions. Write down every place a model output could change a price, a promise, a permission or a payment.
- Write the rule for each one. If it can be a threshold, a table or a checklist, make it one, and make it testable.
- Give the model the reading and drafting. Extracting fields, classifying requests, summarizing and drafting are where variation is cheap.
- Put a deterministic check between the model and the action. Schema validation, limits, an approval step, a log.
- Measure consistency, not only accuracy. Run the same case many times and track how often the answer is the same.
None of this is an argument against using AI. It is an argument for being specific about where its variation costs nothing and where it costs money, a promise or a customer.
Related Reading
- AI and the fractional operator - where an operator with hands-on experience fits when AI changes how work gets done.
- AI readiness assessment guide - how to judge whether a process is ready for AI before building on it.
- The assessment is the product - why a clear diagnosis beats an open-ended retainer.
Sources
- Thinking Machines Lab, "Defeating Nondeterminism in LLM Inference", September 2025. Language-model APIs are not deterministic at temperature 0 because results depend on batch size, which varies with server load; in their test 1,000 identical completions produced 80 unique outputs, and batch-invariant kernels produced identical outputs in all 1,000.
- Atil, B. et al., "Non-Determinism of 'Deterministic' LLM Settings", arXiv 2408.04667. Five LLMs, eight tasks, ten runs: accuracy varied by up to 15% across runs, the best-to-worst gap reached up to 70%, and no model delivered repeatable accuracy across all tasks.
- Zheng, C. et al., "Large Language Models Are Not Robust Multiple Choice Selectors", ICLR 2024 (spotlight). Twenty LLMs, three benchmarks: answers change with option position because of a bias toward particular option IDs.
- Yao, S. et al., Sierra Research, "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024. State-of-the-art function-calling agents succeeded on fewer than 50% of tasks, with pass^8 below 25% in the retail domain.
- Anthropic, "Building effective agents". Defines workflows (predefined code paths) and agents (model-directed processes) and recommends the simplest solution that works.
- Kahneman, D., Rosenfield, A., Gandhi, L. and Blaser, T., "Noise: How to Overcome the High, Hidden Cost of Inconsistent Decision Making", Harvard Business Review, October 2016. The underwriter audit (55% typical difference against about 10% expected) is also described in Kahneman, Sibony and Sunstein, Noise (2021), and in strategy+business, "How noisy is your company?".
- Grove, W., Zald, D., Lebow, B., Snitz, B. and Nelson, C., "Clinical versus mechanical prediction: A meta-analysis", Psychological Assessment 12(1), 2000. 136 studies of human health and behavior; mechanical prediction about 10% more accurate on average.
- Civil Resolution Tribunal of British Columbia, Moffatt v. Air Canada, 2024 BCCRT 149, decided February 14, 2024.
- Gartner, "Gartner Survey Finds 64% of Customers Would Prefer That Companies Didn't Use AI For Customer Service", July 9, 2024. Survey of 5,728 customers conducted in December 2023.
- Scheibehenne, B., Greifeneder, R. and Todd, P., "Can There Ever Be Too Many Options? A Meta-Analytic Review of Choice Overload", Journal of Consumer Research 37(3), 2010. Sixty-three conditions from 50 experiments (N = 5,036); mean effect virtually zero with considerable variance between studies.
- Iyengar, S. and Lepper, M., "When Choice Is Demotivating: Can One Desire Too Much of a Good Thing?", 2000. The jam experiment (24 versus 6 varieties; 3% versus 30% used a voucher) as summarized by AcaWiki.
- Jakob's Law and Nielsen's fourth heuristic: Laws of UX, "Jakob's Law", and Nielsen Norman Group, "Consistency and Standards".
- Findlater, L. and McGrenere, J., "A Comparison of Static, Adaptive, and Adaptable Menus", 2004. The static menu was significantly faster than the adaptive menu, and the majority of participants preferred the adaptable menu.
- Amershi, S. et al., "Guidelines for Human-AI Interaction", CHI 2019, Microsoft Research. Eighteen guidelines validated with 49 practitioners across 20 AI products. See guideline 14, update and adapt cautiously, and guideline 18, notify users about changes.
- The reliability table in section 03 is arithmetic (0.99n and 0.95n), assuming independent steps and a task that fails if any step fails. It is not a measured result.
May Mor
Efficiency Leader. I help operators align their people, systems, and processes so growth scales the business instead of breaking it. M.Sc in AI, 10+ years across fintech, digital banking and adtech. Full bio →