Snap Judgments
Chatbots write and converse. A new kind of AI does something simpler: it answers a narrow question, such as “Is this payment a scam?”, in a split second, says how sure it is, and costs almost nothing. That could change how banks, shops and offices make millions of everyday calls.
Chapters
The audio version is adapted for listening, so some sentences are shorter and some figures are rounded. The voice is generated with OpenAI text-to-speech.
At five past eleven on a Tuesday night, a 68-year-old retired teacher in Pune opens her banking app and tries to send ₹95,000 to a payee she added yesterday. For two days a man claiming to be a police officer has been calling her, and tonight he has kept her on the line for forty minutes. He says her Aadhaar number has turned up in a customs case, and that she must move her savings to a “verification account” until her name is cleared. Indian police have a name for this script: the “digital arrest.” He knows the bank’s rules too. He had her add the payee a day early, so the waiting period for new payees has passed. And he has told her to send the money in parts, each under ₹1 lakh, the amount above which her bank flags transfers to recent payees. The software checks the payee and the amount, finds nothing outside the rules, and is ready to send the money, exactly as it was told.
Every rule has been followed, and the payment is still wrong. The question that matters is one nobody has managed to write as code: does this look like a person being coached by a scammer? The bank can see the signals, but they are scattered: a payment note that reads “customs verification,” a fixed deposit broken an hour earlier, a balance nearly emptied, a transfer at 11 p.m. from a customer who has never banked after eight, and a phone call that her app can see is still running. No single one proves anything, and a rule that acted on any one of them would stop thousands of honest payments. For decades, banks have had two options for cases like this. They tighten the limits and annoy every honest customer, or they accept the losses.
In September, a small company called TypeSafe released a model built for exactly that kind of line. It is called Jev, and you do not ask it to write anything. You give it a situation and a short list of questions with fixed answers: yes or no, one option from a list, or a score from one to five. It returns an answer to each question and a probability, usually in well under a second. TypeSafe’s founder, Diogo Almeida, says the outputs “slot into ordinary software as fuzzy decision rules.” Developers have started calling it a smart if-statement.
The name is a small manifesto. TypeSafe says it named Jev after William Stanley Jevons, the Victorian economist who noticed that more efficient steam engines raised Britain’s demand for coal instead of lowering it. “We expect machine intelligence to follow a similar path,” the company wrote. If judgment becomes nearly free, the bet goes, people will use far more of it.
Two weeks later, OpenAI made a similar bet. At its DevDay conference on Sept. 29, the company announced a Decisions API, in limited preview, that lets developers give its Luna model a fixed set of options to choose from. “By focusing the model on that choice, we can make it extremely fast while keeping capabilities like image understanding, broad language support, and safety protections,” Sam Altman said, according to TechCrunch. OpenAI has not yet published prices or response times. A category is forming. In this essay we call these tools judgment models, and the compact, low-cost versions small judgment models, or SJMs. Other names are in use too; we explain our choice near the end.
Thinking, fast and cheap
TypeSafe calls Jev a “System One” model, after the psychologist Daniel Kahneman’s Thinking, Fast and Slow. In Kahneman’s account the mind runs two processes: a fast, automatic one that recognizes a face or senses danger, and a slow, effortful one that does long division. Large language models such as Claude or GPT, asked to reason step by step, resemble the slow system. They are very good at it, and for many small decisions they are also slow and expensive.
TypeSafe’s claims are bold. It says Jev answers in 70 to 500 milliseconds, against “3 to 329 seconds” for frontier models on its workflows, and it charges $0.042 per million input tokens, with output “too cheap to meter.” It also says the probabilities carry meaning: “higher confidence means higher accuracy.” These are the company’s own numbers, measured on workflows it designed, with reference answers produced by other AI models. A DataCamp review of those results found Jev roughly level on accuracy with a mid-tier frontier model at about one-seventieth of the cost per case, and still behind the strongest model in the comparison. The reviewer called the result “promising rather than settled.” The consultancy innFactory went further, noting that no independent benchmarks exist yet and that TypeSafe cannot prove its price is not subsidized.
That skepticism is fair. The idea underneath Jev stands on its own, though, and it is worth understanding before any vendor’s numbers are settled.
Anatomy of a smart if-statement
Look closely at almost any business decision and it breaks into smaller ones. “Should we enter Mumbai?” becomes: Is demand proven? Do the unit economics work? How strong is the competition? Which approach fits the evidence? A few of these can be computed. Most need judgment.
The design Jev encourages splits the work three ways. Arithmetic stays in code, where it is exact. Judgment goes to the model in bounded form, with answers the developer defines in advance. The decision rules, meaning the thresholds that turn judgments into action, go back into code, where managers can read it, argue about it and change it.
The split also changes who controls the decision. Ask an AI system “what should we do?” and the rules disappear inside a prompt. Ask it “is demand proven, and how sure are you?” and the rules stay on the table. A finance director can approve claims automatically below a suspicion score of 0.2, send anything above 0.8 straight to audit, and move those numbers next quarter without touching the model.
Three ways to run the same process
Take any process that ends in a decision: approving expense claims, triaging support tickets, screening suppliers. Today most companies run it the old way, with hard-coded rules for the easy cases and a queue of human reviewers for everything else. The rules are cheap and predictable but brittle. The queue is accurate when people have time, and slow when they do not.
The current wave of agentic AI proposes a second way: give the whole task to a large language model, with tools, and let it read, reason, decide and act. Agents are flexible. They can handle a case nobody anticipated and explain themselves in fluent prose. They are also slow and expensive per decision, they can vary from one run to the next, and the rules they follow live in a prompt that few people read.
Judgment models suggest a third way. Break the process into bounded judgments, let a fast model make them, keep the decision rules in code, and call a large model or a person only when the small one is unsure or the case is new. The approaches combine well. An agent can plan the steps, while judgment models make the many small calls inside each step and decide when to stop and ask a person.
| Rules plus people | An LLM agent does the task | Judgment models plus rules | |
|---|---|---|---|
| Speed per decision | Instant for rules; hours or days in the queue | Seconds to minutes | Under a second |
| Cost per decision | Low for rules; high for reviewers | Cents to dollars | Fractions of a cent |
| Handles messy input | Only people do | Yes | Yes, within fixed answers |
| Handles a truly new case | People do | Often | Hands it to a person or a large model |
| Where the rules live | Code and the reviewer’s head | A prompt | Code that managers can read |
| Audit trail | Partial | Long transcripts | Inputs, answers, probabilities and the rule |
| Typical failure | Backlogs and inconsistency | Confident, hard-to-trace errors | Badly framed questions or thresholds |
What our classroom test found
To see how this works in practice, we built a small benchmark for an MBA classroom. We wrote 18 fictional case files covering market entry, investment committees, product launches, supplier bids, consulting diagnoses and a simulated credit committee. Each runs to about a thousand words across seven documents: memos, tables, interview notes and news clippings, with a misleading quote from an insider and at least one figure that is corrected later in the file. We sent the same bounded questions about each case to Jev and to Anthropic’s Claude Opus 4.8, which had to return a probability for every option in a fixed format. The same decision rules turned both sets of answers into a decision. Each case ran five times on each model.
| Decision | One of the cases | Examples of the bounded questions | Possible actions |
|---|---|---|---|
| Market entry | A grocery-delivery startup weighing a move into Mumbai | Is demand in the new city proven? (yes/no) · How strong is the competition? (1 to 5) · Which approach fits the evidence? (pick one) | Expand · pilot · fix home markets |
| Investment committee | A payables-software startup asking for Series A money | Do customers pay, stay and use more? (yes/no) · How hard is it to copy? (1 to 5) · How much does it depend on one customer? (1 to 5) | Partner meeting · more diligence · pass |
| Product launch | A cold-brew snack bar after a 16-week test market | Is the sales forecast realistic? (yes/no) · What is the biggest risk? (pick one) | National launch · regional pilot · delay |
| Supplier selection | Three bids to supply 400,000 brake discs a year | Is the quality proven? (yes/no) · How reliable are deliveries? (1 to 5) · Are the contract terms acceptable? (yes/no) | Award · negotiate · reject |
| Consulting diagnosis | A restaurant chain whose profit fell 18 percent | What mainly caused the fall? (pick one) · Are fewer customers coming? (yes/no) · Did costs go up? (yes/no) | Study demand · costs · prices · all |
| Credit committee (fictional) | A gear maker asking for a ₹3 crore loan | Is the growth plan realistic? (yes/no) · Does repayment depend on optimistic assumptions? (1 to 5) | Approve · approve with conditions · review · decline |
Based on … measured runs on …. Jev time is the full round trip from a laptop; the Opus time leaves out start-up overhead, which favors Opus. Costs are measured tokens times published prices. The case author wrote the reference answers, and faculty have not reviewed them yet.
The headline gaps were large. Jev’s typical decision took …, against … for Opus 4.8, and every Jev call in the benchmark finished in under …. Opus was usually quick for a frontier model: 95 percent of its calls finished within …. Projected to 1,000 decisions, Jev would cost … and Opus … at published prices. Jev also read less: about … input tokens per case, against … for Opus, which also produced about … output tokens of its own.
Accuracy was closer, and in this small test it did not favor the bigger model. Jev’s final decision matched our reference in … of runs, Opus’s in …, and the two reached the same decision in … cases. With 18 cases and a single author writing the references, none of this proves that one model judges better than the other. It does show a fast, inexpensive model holding its own on long, messy documents.
The disagreements were the most useful part. In one case, a pharmacy chain planning to enter Pune, both models recommended a pilot where our reference said to expand at full scale. The only evidence about Pune was a survey, and on reflection the models had a point. In another, a grocery-delivery startup weighing Mumbai, Opus preferred to fix the home markets first, a defensible reading of an eight percent negative margin.
The most instructive mistake was ours. In an earlier run, both models and our reference disagreed on that same Mumbai case for a reason that had nothing to do with AI. Our decision rules allowed a pilot only once demand in the new city was already proven. Uncertain demand is the reason companies run pilots, so the rule was backward. Because the rule lived in code, the mistake took one line to find and one minute to fix, and we could test the fix against every stored answer without paying for a single new model call. If we had asked a model to “solve the case,” the same flaw would have sat unnoticed inside a paragraph of confident prose.
Is this really a big deal?
Many people hear the pitch and shrug, with reason. Software that sorts things into categories is old. Bayesian spam filters spread in the early 2000s. Banks have scored credit with statistical models for decades, and card networks flag fraud in milliseconds. Content moderation, sentiment analysis and document routing all run on classifiers. In that sense a judgment model is a classifier with a new name.
What changed is the cost of getting one. A traditional classifier needed thousands of labeled examples, a data science team, a separate model for each question and steady maintenance as the world drifted. Most companies built a handful for their biggest problems and handled everything else with rules and people. A judgment model asks for none of that. You write the question and the possible answers in plain English, send the documents, and get a probability back. There is no training step, one interface covers any question, and several questions about the same case travel in one call. In our test it read thousand-word case files with tables and contradictory memos, the kind of input that classic classifiers handle badly.
The timing matters as much as the technology. Companies are now trying to use large language models and agents for exactly these small decisions: tagging tickets, screening invoices, checking whether an agent may take its next step. A frontier model can do that, as Opus 4.8 did in our test, but at a cost and speed that make it uneconomical to run on every email and every row. Judgment models offer a middle path between brittle rules and an expensive general model, and they arrive while budgets for AI are being spent and the habits are still forming.
A useful comparison is cloud storage. Nothing about storing files was new when online storage became cheap and simple to call from code. What changed was who could use it and how often. If judgment models follow that path, the big effect will come from the many small decisions that were never worth automating before, more than from any single impressive answer.
A large model in a smaller box?
Partly, and that matters. Small models usually learn from big ones, a technique called distillation that Geoffrey Hinton and colleagues formalized in 2015. Meta trained its smallest Llama models on the outputs of larger siblings. Microsoft trained one of its Phi reasoning models on problems generated by DeepSeek’s R1. TypeSafe describes a new architecture and a training method it calls “Reinforcement Learning for Calibrated Decisions,” but has not said what Jev is built on. A reviewer at DigitalOcean wrote that the company was “just as quiet about how Jev got built as it was about what it’s actually running on.”
So the quality of a Jev-like tool is, to a large degree, inherited. A decision model is unlikely to judge much better than the teachers and data it learned from. Distillation is also contested. In February, Anthropic said it had detected about 24,000 fraudulent accounts used to harvest more than 16 million exchanges from its models, while acknowledging that distillation in general is “a widely used and legitimate training method.”
The large-model companies are moving too. OpenAI’s structured outputs, introduced in 2024, made frontier models follow a required JSON format reliably. GPT-5 shipped with a router that decides when to think hard and when to answer fast. Anthropic lets developers set how much effort Claude spends. And with the Decisions API, OpenAI now offers the smart if-statement as a product of its own.
Jevons, again
Will decision models eat into large-model usage? Some calls will move. Classifying a ticket, scoring a lead or flagging an invoice can go to something faster and cheaper than a frontier model. The history behind Jev’s name suggests the larger effect runs the other way.
The price of a given level of AI capability has been falling by roughly an order of magnitude a year, according to an analysis by the venture firm Andreessen Horowitz, and by anywhere from 9 to 900 times a year depending on the task, according to the research group Epoch AI. Usage has climbed even faster. When the Chinese lab DeepSeek showed in early 2025 that strong models could be trained cheaply, Microsoft’s chief executive, Satya Nadella, responded: “Jevons paradox strikes again!”
Decision models push that logic further. At a fraction of a cent per thousand judgments, it becomes rational to put a judgment on every incoming email, every spreadsheet row and every step an AI agent takes. Most of those judgments were never made before, because they were not worth a person’s time or a frontier model’s fee. Cheap decisions also create work for large models: writing and testing the decision rules, explaining edge cases, generating training data and handling the questions the small model is unsure about. The likely result is a division of labor, with fast judgments at enormous volume and slow reasoning where it earns its cost.
The router in the middle
One natural job for a decision model is choosing which other model to use. Researchers have shown the potential for years. Stanford’s FrugalGPT project reported matching the best single model with cost reductions of up to 98 percent by sending easy questions to cheaper models first. LMSYS’s RouteLLM reported cost cuts of more than 85 percent on one benchmark while keeping 95 percent of GPT-4’s quality. Commercial routers claim smaller savings, typically 20 to 30 percent, which still adds up at enterprise scale. TechCrunch reported that OpenAI sees its own decision model partly as a way to keep its swarms of AI agents in check.
A decision model with calibrated probabilities makes a natural switchboard: send the request to a small model when it is confident, to a large model when it is not, and to a person when the stakes or the uncertainty are high. AI agents face the same choice dozens of times per task, and each choice is a bounded judgment. Making it in a quarter of a second instead of ten seconds changes what agents can practically do.
Small enough for your pocket
Apple now gives developers an on-device model of about three billion parameters in “as few as three lines” of code, and Google offers Gemini Nano on Android phones with “no additional cost incurred for each API call.” Researchers at NVIDIA argued last year that small language models are “sufficiently powerful, inherently more suitable, and necessarily more economical” for most agent tasks. They estimated that serving a 7-billion-parameter model is 10 to 30 times cheaper than serving one with 70 to 175 billion, and that small models could handle 40 to 70 percent of the model calls in the agent systems they studied.
Jev itself runs in the cloud. The work it does, narrow judgments with short outputs in strict formats, is exactly the kind that fits on a phone, and this is where small judgment models, or SJMs, should appear first. A field-sales app could triage leads offline, a banking app could flag an odd transfer before the network connects, and a clinic tablet could route an intake form without sending patient text to a distant server.
The decision-driven organization
Gartner predicts that by 2027, half of business decisions will be “augmented or automated by AI agents.” If that happens, the important management question will be how decisions are designed, more than which model to buy.
With smart if-statements, a decision has four parts anyone can point to: the data that went in, the bounded judgments, the decision rules and the log. Thresholds become management levers. A risk committee that wants fewer false alarms raises a number; a growth team that wants more experiments lowers one. Because every judgment comes with a probability, an organization can check its machines the way it should check its people: were the 80 percent calls right about 80 percent of the time? A 2022 Anthropic study found that large models can be well calibrated when questions are posed in the right format, and a 2023 study found that models asked to state their confidence in words tend to be overconfident. The format matters.
Regulators are moving in the same direction. Europe’s AI Act requires that people overseeing high-risk systems remain aware of “automation bias” and be able to “disregard, override or reverse the output.” After this summer’s amendments, the rules for stand-alone high-risk uses such as credit scoring and hiring are set to apply from December 2027. A system in which the model makes narrow, logged judgments and the decision rules sit in readable code is far easier to explain to an auditor than one in which a single model does everything.
How this will rewire decisions and processes
If judgment models become as cheap and common as their makers expect, the effects will reach well beyond the IT department. Seven predictions:
- Review queues will shrink and change shape. Machines will settle the clear cases at both ends. People will spend their time on the uncertain middle, where a probability of 0.4 to 0.7 says a human should look.
- Decision rules will become products. Thresholds and rules will be versioned, tested on past cases and changed every week, with a named business owner. “Rule owner” will become a real job title in operations and risk teams.
- Process maps will be redrawn around decisions. Companies will list the judgments their processes depend on, the way they once listed their applications, and ask which ones a model can make, which ones a person must make, and which ones nobody makes today.
- Periodic reviews will become continuous. When a judgment costs a fraction of a cent, a credit line, a supplier’s risk or a customer’s churn risk can be re-scored every day instead of every quarter.
- Every agent action will pass through a gate. Before an AI agent sends money, deletes a file or emails a customer, a fast judgment model will check whether the step is allowed, and the decision rules will say whether a person must approve it.
- Calibration reviews will become a management ritual. Like a monthly financial close, teams will compare what their models said with what happened and adjust thresholds, using the decision log as the evidence.
- Decision design will become a core skill. Managers will need to frame good bounded questions, choose sensible thresholds and decide what a person must see. Business schools will teach it alongside spreadsheets.
Some organizations will automate too much, too fast. The ones that benefit will be those that treat every automated judgment as a decision someone owns, with a rule they can read and a log they can check.
A name for the category: judgment models and SJMs
The field does not have a settled name yet. These are the terms in use or on offer:
- System 1 models. TypeSafe’s term, after Kahneman’s fast mode of thinking. Apt, but it belongs to one vendor.
- Decision models, echoing OpenAI’s Decisions API. Clear, but the phrase already means something else in operations research and in business-rules software.
- Smart switches or smart if-statements. Good everyday shorthand for what the tools do inside code.
- Typed decision models. Precise about the fixed output types, and mainly of interest to developers.
- Judgment models. The term we use for the category.
- Small judgment models (SJMs). A term coined by AI Profs for compact, low-cost versions that can run at very large volume or on devices, by analogy with small language models (SLMs). TypeSafe has not disclosed Jev’s size, so we do not assume that every judgment model is small.
We define a judgment model as an AI model that takes a situation and a fixed set of questions, and returns a typed answer and a calibrated probability for each, quickly and cheaply enough to sit inside ordinary software. Jev and OpenAI’s Decisions API are the first products built for the job. Five traits define the category:
- Bounded output. The developer defines every possible answer in advance; the model never writes free text.
- A probability with every answer, and a claim that the probability is calibrated, which buyers should test.
- Speed and price fit for ordinary code: well under a second and well under a cent per decision.
- Many questions about one situation in a single call.
- Decision rules kept outside the model, where people can read and change them.
What could go wrong
The headline numbers for Jev come from the vendor, and its price may not last. OpenAI has not yet published prices or response times for its version. A reviewer quoted by DigitalOcean warned that Jev “performs far less reliably when the question hides several judgments,” which is exactly the kind of question tired product teams will write. Calibration claims have not been independently verified. Cheap, confident scores invite over-trust, the automation bias the law warns about. Cheaper judgment may also raise total computing and energy use, which is the Jevons paradox at work. And Anthropic’s data on business use suggests that buyers care more about capability than price, which could limit how far the cheap tier spreads.
What comes next for judgment models
Every major vendor will offer one. OpenAI has already joined TypeSafe. Google, Anthropic, Microsoft and the cloud providers each have small, fast models today, and we expect them to package those models in this shape, as a separate product or as a mode of an existing one.
Open-source versions will follow. Small open-weight models, combined with existing open-source tools that force a model’s output into a fixed format, already make a basic judgment model possible on a company’s own servers or on a phone. Expect projects that are tuned and calibrated specifically for decisions.
Domain-specific versions will matter most. Credit, insurance claims, clinical intake, legal triage, customs and content moderation each have their own vocabulary and their own regulators. Models trained on domain data, with documented behavior, will have an advantage in those markets.
Judgments will move beyond text. Jev reads text only: the situation arrives as written documents and data. Many everyday judgments start elsewhere. An insurer wants to know whether a photo shows hail damage or old rust. A bank wants to know whether the voice on a call sounds coached. A factory wants to know whether a few seconds of video show a safety lapse. Large models already read images, audio and video, so judgment models that take the same inputs and still return a typed answer with a probability are a natural next step. They will cost more and run slower than text-only versions at first, and audio and video raise sharper questions about consent and privacy.
Benchmarks will have to catch up. The category needs neutral tests that measure what buyers care about: agreement with expert references, calibration, 95th-percentile response time, cost per 1,000 decisions, stability across repeated runs, and how well a model copes with long, messy input and with questions that hide several judgments. Our classroom harness is a very small version of such a test.
Buyers will ask what is inside. Because a judgment model inherits much of its quality from its teachers, customers will want model cards that say which models and data shaped it, how it was calibrated and how it is updated. A model name that silently changes underneath an application is a governance risk.
The interface may become a standard. “A situation, typed questions, answers with probabilities” is simple enough to standardize, much as databases settled on SQL. A shared format would let companies switch vendors, compare them on the same cases and keep their decision rules unchanged.
Back in Pune, the payment has not gone yet. Under the old rules it would have. With a smart if-statement, the bank’s software asks one plain question about everything it sees and gets back “yes, 0.91.” The bank’s risk committee has signed a rule for that: above 0.8, pause the transfer whatever the amount, and call the customer. Someone from the bank calls her before the money moves. For the next customer, a son paying his mother’s hospital bill at 0.12, the payment goes through and nobody is bothered. AI judgment day is here. It arrives quietly, as a pause on a payment in Pune, five minutes past eleven at night.
Sources
- TypeSafe, “Introducing System One Models & Jev,” Sep 15, 2026
- TechCrunch, “OpenAI’s Jev clone could help the frontier lab stop its swarming agents,” Sep 30, 2026
- Axios, “The 5 biggest announcements from OpenAI’s blockbuster AI conference,” Sep 29, 2026
- DataCamp, review of System One models and Jev, Sep 16, 2026
- innFactory on Jev, Sep 21, 2026
- DigitalOcean, “What is Jev?”, updated Sep 24, 2026
- Hinton, Vinyals, Dean, “Distilling the Knowledge in a Neural Network,” 2015
- Meta, Llama 3.2 small models, Sep 2024
- TechCrunch on Microsoft Phi-4 reasoning models, Apr 30, 2025
- Anthropic, “Detecting and preventing distillation attacks,” Feb 23, 2026
- OpenAI, Structured Outputs, Aug 6, 2024
- OpenAI, Introducing GPT-5, Aug 7, 2025
- Anthropic, Claude Opus 4.5 and the effort parameter, Nov 24, 2025
- Andreessen Horowitz, “LLMflation,” Nov 12, 2024
- Epoch AI, LLM inference price trends, Mar 12, 2025
- Fortune on Satya Nadella and the Jevons paradox, Jan 27, 2025
- Shacknews, Google’s 3.2 quadrillion monthly tokens, May 2026
- The Decoder on Google’s monthly token counts
- Chen, Zaharia, Zou, “FrugalGPT,” May 2023
- LMSYS, RouteLLM, Jul 1, 2024
- The Pragmatic Engineer on smart model routing, Jul 2, 2026
- TechCrunch on Apple’s Foundation Models framework, Jun 9, 2025
- Android Developers, ML Kit GenAI APIs with Gemini Nano, May 20, 2025
- NVIDIA, “Small Language Models are the Future of Agentic AI,” Jun 2025
- Gartner, top data and analytics predictions, Jun 17, 2025
- Kadavath et al., “Language Models (Mostly) Know What They Know,” 2022
- Xiong et al., on verbalized confidence in LLMs, 2023
- EU AI Act, Article 14: human oversight
- Hunton, EU Digital Omnibus on AI enters into force, Jul 2026
- Anthropic Economic Index, Sep 15, 2025
Vendor claims are attributed to the vendor. Numbers about our classroom test come from this project’s own result files and are calculated each time the page loads. Diagrams marked “Illustration” do not show measured data. “Judgment models” and “small judgment models (SJMs)” are the names we use for the category.