Skip to content
AI News HubLIVE
In-site rewrite6 min read

A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model

Summary

This tutorial provides a complete coding guide to TypeSafe AI's Jev, a System One model designed for non-text, structured judgments. It covers installing the official Python SDK, using primitive question types (Choice, Score, Noul), implementing speculative fan-out, confidence-gated routing, and building async production workflows The post A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model appeared first on MarkTechPost.

SourceMarkTechPostAuthor: Asif Razzaq
A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model
Report an error

The correction channel is not available yet. You can copy the article reference below for later.

Correction instructions
Read article

In this tutorial, we work with Jev, TypeSafe AI’s first System One model, which does not generate text at all: we send it a piece of program state and a set of typed questions, and it returns choices, scores, and yes/no probabilities that our code can branch on directly. We install the official Python SDK, make a first call that uses all three question primitives at once, and look at how the shape of the state changes what the model can know. We then recompute the published confidence statistic from the returned probabilities, measure what batching ten questions into one call buys over ten separate calls, and build the patterns the API is designed for: confidence-gated routing, composite scoring with the weights kept in code, typed function calling, and counting done the way the model can actually do it. We close with the production shape: Pydantic response models, an async client fanned out with asyncio, retry policies, typed errors, and a running ledger that prices the whole notebook. Copy CodeCopiedUse a different Browser import os import sys import json import time import asyncio import traceback import subprocess from getpass import getpass RESULTS = {} LEDGER = {"calls": 0, "input_tokens": 0, "output_tokens": 0} USD_PER_MILLION_INPUT_TOKENS = 0.042 # Jev list price; output tokens are free def banner(title): print("\n" + "=" * 78) print(title) print("=" * 78) def section(name): def wrap(fn): def run(*a, kw): banner(name) try: out = fn(*a, kw) RESULTS[name] = out if isinstance(out, str) else "ok" return out except Exception as e: RESULTS[name] = f"SKIPPED / FAILED -> {type(e).name}: {e}" print(f"\n[!] {name} did not complete: {type(e).name}: {e}") traceback.print_exc(limit=3) return None return run return wrap banner("0. Install the SDK, load the API key, list the models") subprocess.run([sys.executable, "-m", "pip", "install", "-q", "typesafe-sdk==0.7.0"], check=True) import typesafe_sdk from typesafe_sdk import Choice, Noul, Score, TypeSafeClient def load_api_key(): key = os.environ.get("TYPESAFE_API_KEY", "").strip() if not key: try: from google.colab import userdata # Colab: key stored under the Secrets tab key = (userdata.get("TYPESAFE_API_KEY") or "").strip() except Exception: key = "" return key or getpass("TypeSafe API key (console.typesafe.ai/keys): ").strip() os.environ["TYPESAFE_API_KEY"] = load_api_key() client = TypeSafeClient() # reads TYPESAFE_API_KEY, defaults to jev-latest print(f" typesafe-sdk {typesafe_sdk.version} | Python {sys.version.split()[0]}") print(" models available to this key:") for m in client.models.list().models: print(f" {m.name: {dept.choice!r} confidence {dept.confidence:.3f}") print(f" probabilities {({k: round(v, 3) for k, v in dept.probabilities.items()})}") fr = response.scores["frustration"] print(f" frustration -> score {fr.score:.3f} on 0..{len(fr.legend) - 1} confidence {fr.confidence:.3f}") for level, text in fr.legend.items(): print(f" {level}: p={fr.probabilities[level]:.3f} {text}") print(f" refund_requested -> noul {response.nouls['refund_requested'].noul:.3f}") print(f" policy_supports -> noul {response.nouls['policy_supports'].noul:.3f}") print(f"\n answered by {response.model} in {ms:.0f} ms " f"input tokens {response.usage.input_tokens}, output tokens {response.usage.output_tokens}") return f"{dept.choice}, frustration {fr.score:.2f}, refund {response.nouls['refund_requested'].noul:.2f}" three_primitives() A System One request has two parts: state, which is any text, JSON object or array describing the situation, and a dictionary of named questions. Choice selects one label from the criteria we define and returns a probability for every label; Score places the state on an ordered rubric and returns the probability-weighted level, so it can land between two levels; Noul returns a single probability that a statement is true. The question names are ours and never reach the model, which is why the instructions carry the full meaning and can point at nested fields with backticked paths. All four questions are evaluated in one request, in parallel and in isolation from one another, and the response reports the pinned model version that answered and the tokens it billed. Copy CodeCopiedUse a different Browser @section("2. State is program state: the same question over a string and over named fields") def state_shapes(): question = {"eligible": Noul( instructions="The customer is eligible for a refund under the company's written policy", criteria={"true": "A policy is present and it covers the customer's situation", "false": "No policy is given, or the policy does not cover the situation"}, )} bare = "I was charged twice for order A-104. Please refund the duplicate." as_list = [m["text"] for m in TICKET["ticket"]["messages"]] shapes = [("string: the message only", bare), ("array : the conversation", as_list), ("object: ticket + order + policy", TICKET)] print(f" {'state shape':6s} input tokens ms") seen = {} for label, state in shapes: response, ms = ask(state, question) seen[label] = response.nouls["eligible"].noul print(f" {label:8s} {'recomputed':>11s} " f"{'score':>6s} {'sum(level*p)':>13s} {'API conf':>9s}") worst = 1.0 for label, text in messages.items(): response, _ = ask(text, {"tone": tone, "urgency": urgency}) t, u = response.choices["tone"], response.scores["urgency"] expected = sum(level * p for level, p in u.probabilities.items()) print(f" {label:10s} {'own call':>10s}") for name, q in FANOUT.items(): single, ms = ask({"postmortem": POSTMORTEM}, {name: q}) seq_ms, seq_tokens = seq_ms + ms, seq_tokens + single.usage.input_tokens a, b = value_of(batched.answers[name]), value_of(single.answers[name]) same = a == b if isinstance(a, str) else abs(a - b) 10s}") if isinstance(a, str) else (lambda v: f"{v:10.3f}") print(f" {name: {seq_ms / batched_ms:.1f}x faster and {seq_tokens / batched_tokens:.1f}x fewer tokens; " f"{agree}/{len(FANOUT)} answers agree, because questions never see each other") return f"{seq_ms / batched_ms:.1f}x faster, {seq_tokens / batched_tokens:.1f}x cheaper, {agree}/{len(FANOUT)} agree" fan_out() Because questions in a request cannot see each other, we can ask everything we might need up front, including questions that only matter on one branch, and read only the relevant answers afterwards. We put ten questions about an incident postmortem, two Choices, two Scores and six Nouls, into one call, then ask each of them again in a call of its own, and compare wall time, input tokens and the answers. The state is sent once instead of ten times, which is where both the latency and the token savings come from, and the agreement column checks the isolation claim directly: a question should receive the same answer whether or not it travels with others. Copy CodeCopiedUse a different Browser INTENT = Choice( instructions="What the user wants the banking assistant to do", criteria={"check_balance": "See a balance or recent transactions", "approve_transfer": "Send or approve a transfer of money", "dispute_charge": "Contest a charge they do not recognise", "close_account": "Close the account permanently", "other": "Anything else, or not clear enough to act on"}, ) STAKES = {"check_balance": 0.50, "dispute_charge": 0.70, "approve_transfer": 0.85, "close_account": 0.90} def route(answer): if answer.choice == "other" or answer.confidence human" bar = STAKES[answer.choice] return f"-> run {answer.choice}" if answer.confidence >= bar else f"-> confirm first (needs {bar:.2f})" @section("5. Confidence-gated routing: the bar rises with the stakes") def gated_routing(): inbox = ["how much is in my checking account", "send 2,000 to my landlord like last month", "i guess maybe move some money around? not sure", "there's a 89.99 charge from a gym i never joined", "shut everything down, i'm done with this bank", "what's the weather like in lisbon"] print(f" {'message':5s} decision") acted = 0 for text in inbox: response, _ = ask(text, {"intent": INTENT}) a = response.choices["intent"] decision = route(a) acted += decision.startswith("-> run") print(f" {text[:50]:15s}" for d in DIMENSIONS) + " (each normalised to 0..1)") for name, row in table.items(): print(f" {name: ".join(f"{n} {sum(w[d] * table[n][d] for d in w):.2f}" for n in ranked)) print("\n Two rankings, four model calls: changing the weights re-ran no inference.") return ", ".join(f"{role}: {who}" for role, who in winners.items()) composite_scoring() Composite scoring keeps the model’s job narrow and the policy explicit. For each candidate we ask four Score questions, each describing concrete situations rather than degrees, normalise every score by its top level, and store the resulting table. The ranking is then plain arithmetic: one weight vector for a senior individual contributor, another for a team lead. Because the judgments are stored separately from the weights, changing what we value instantly re-ranks the candidates and requires no inference. You can trace every position in the ranking back to the dimension that produced it. Copy CodeCopiedUse a different Browser ROOMS = {"living_room": None, "bedroom": None, "kitchen": None, "office": None} def set_lights(room, state): return f"lights in {room} -> {state}" def set_thermostat(room, mode): return f"thermostat in {room} -> {mode}" def play_music(room, genre): return f"playing {genre} in {room}" TOOLS = {"set_lights": (set_lights, "state"), "set_thermostat": (set_thermostat, "mode"), "play_music": (play_music, "genre")} CALL_SPEC = { "tool": Choice(instructions="Which smart-home function the command asks for", criteria={"set_lights": "Turn lights on, off, or dim them", "set_thermostat": "Make a room warmer, cooler, or set eco mode", "play_music": "Play music or audio", "none": "Not a smart-home command this system supports"}), "room": Choice(instructions="Which room the command refers to", criteria=ROOMS), "state": Choice(instructions="If this is a lights command: the requested light state", criteria={"on": None, "off": None, "dim": None}), "mode": Choice(instructions="If this is a thermostat command: the requested mode", criteria={"heat": "Warmer", "cool": "Cooler", "eco": "Energy saving"}), "genre": Choice(instructions="If this is a music command: the requested genre", criteria={"jazz": None, "classical": None, "rock": None, "ambient": None}), } @section("7. Typed function calling, and counting the way Jev can do it") def function_calling(): commands = ["it's freezing in the office, warm it up", "kill the lights in the bedroom", "put on something mellow and jazzy in the kitchen", "order me a pizza"] dispatched = 0 for text in commands: response, ms = ask(text, CALL_SPEC) # every argument asked speculatively, one call c = response.choices tool = c["tool"].choice if tool == "none": print(f" {text!r: no tool (confidence {c['tool'].confidence:.2f})") continue fn, arg = TOOLS[tool] weakest = min(c["tool"].confidence, c["room"].confidence, c[arg].confidence) print(f" {text!r: {tool}(room={c['room'].choice!r}, {arg}={c[arg].choice!r}) " f"weakest judgment {weakest:.2f}, {ms:.0f} ms") print(f" {'': 0.5 for p in probs) print(f" fruits counted: {count} of {len(basket)}") return f"{dispatched}/{len(commands)} commands dispatched from typed answers; counted {count} fruits" function_calling() Function calling becomes a set of closed-set questions: one Choice selects the tool, including an explicit none option for commands we do not support, and one Choice per argument is asked speculatively in the same request. The code reads only the arguments that belong to the selected tool, reports the weakest judgment as the confidence of the whole call, and then executes an ordinary Python function with validated, enumerated values. The second half applies a documented workaround: Jev does not count reliably inside a single question, so we ask one Noul per item in one request and do the sum in code. Copy CodeCopiedUse a different Browser import [truncated for AI cost control]

Key points and analysis

Article intelligence

EngineersIntermediate

Key points

  • AI generation is temporarily unavailable; this entry was preserved with deterministic fallback metadata.
  • This tutorial provides a complete coding guide to TypeSafe AI's Jev, a System One model designed for non-text, structured judgments. It covers installing the official Python SDK,…

Highlights and analysis are generated automatically and may contain errors. Check the original source.