跳到主要內容
AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:This tutorial provides a complete coding guide to TypeSafe AI's Jev, a System One model designed for non-text, structured judgments. It covers installing the official Python SDK, using primitive question types (Choice, Score, Noul), implementing speculative fan-out, confidence-gated routing, and building async production workflows The post A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model appeared first on MarkTechPost.

來源MarkTechPost作者: Asif Razzaq
待翻譯:A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model
回報錯誤

更正管道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

In this tutorial, we work with Jev, TypeSafe AI’s first System One model, which does not generate text at all: we send it a piece of program state and a set of typed questions, and it returns choices, scores, and yes/no probabilities that our code can branch on directly. We install the official Python SDK, make a first call that uses all three question primitives at once, and look at how the shape of the state changes what the model can know. We then recompute the published confidence statistic from the returned probabilities, measure what batching ten questions into one call buys over ten separate calls, and build the patterns the API is designed for: confidence-gated routing, composite scoring with the weights kept in code, typed function calling, and counting done the way the model can actually do it. We close with the production shape: Pydantic response models, an async client fanned out with asyncio, retry policies, typed errors, and a running ledger that prices the whole notebook. Copy CodeCopiedUse a different Browser import os import sys import json import time import asyncio import traceback import subprocess from getpass import getpass RESULTS = {} LEDGER = {"calls": 0, "input_tokens": 0, "output_tokens": 0} USD_PER_MILLION_INPUT_TOKENS = 0.042 # Jev list price; output tokens are free def banner(title): print("\n" + "=" * 78) print(title) print("=" * 78) def section(name): def wrap(fn): def run(*a, kw): banner(name) try: out = fn(*a, kw) RESULTS[name] = out if isinstance(out, str) else "ok" return out except Exception as e: RESULTS[name] = f"SKIPPED / FAILED -> {type(e).name}: {e}" print(f"\n[!] {name} did not complete: {type(e).name}: {e}") traceback.print_exc(limit=3) return None return run return wrap banner("0. Install the SDK, load the API key, list the models") subprocess.run([sys.executable, "-m", "pip", "install", "-q", "typesafe-sdk==0.7.0"], check=True) import typesafe_sdk from typesafe_sdk import Choice, Noul, Score, TypeSafeClient def load_api_key(): key = os.environ.get("TYPESAFE_API_KEY", "").strip() if not key: try: from google.colab import userdata # Colab: key stored under the Secrets tab key = (userdata.get("TYPESAFE_API_KEY") or "").strip() except Exception: key = "" return key or getpass("TypeSafe API key (console.typesafe.ai/keys): ").strip() os.environ["TYPESAFE_API_KEY"] = load_api_key() client = TypeSafeClient() # reads TYPESAFE_API_KEY, defaults to jev-latest print(f" typesafe-sdk {typesafe_sdk.version} | Python {sys.version.split()[0]}") print(" models available to this key:") for m in client.models.list().models: print(f" {m.name: {dept.choice!r} confidence {dept.confidence:.3f}") print(f" probabilities {({k: round(v, 3) for k, v in dept.probabilities.items()})}") fr = response.scores["frustration"] print(f" frustration -> score {fr.score:.3f} on 0..{len(fr.legend) - 1} confidence {fr.confidence:.3f}") for level, text in fr.legend.items(): print(f" {level}: p={fr.probabilities[level]:.3f} {text}") print(f" refund_requested -> noul {response.nouls['refund_requested'].noul:.3f}") print(f" policy_supports -> noul {response.nouls['policy_supports'].noul:.3f}") print(f"\n answered by {response.model} in {ms:.0f} ms " f"input tokens {response.usage.input_tokens}, output tokens {response.usage.output_tokens}") return f"{dept.choice}, frustration {fr.score:.2f}, refund {response.nouls['refund_requested'].noul:.2f}" three_primitives() A System One request has two parts: state, which is any text, JSON object or array describing the situation, and a dictionary of named questions. Choice selects one label from the criteria we define and returns a probability for every label; Score places the state on an ordered rubric and returns the probability-weighted level, so it can land between two levels; Noul returns a single probability that a statement is true. The question names are ours and never reach the model, which is why the instructions carry the full meaning and can point at nested fields with backticked paths. All four questions are evaluated in one request, in parallel and in isolation from one another, and the response reports the pinned model version that answered and the tokens it billed. Copy CodeCopiedUse a different Browser @section("2. State is program state: the same question over a string and over named fields") def state_shapes(): question = {"eligible": Noul( instructions="The customer is eligible for a refund under the company's written policy", criteria={"true": "A policy is present and it covers the customer's situation", "false": "No policy is given, or the policy does not cover the situation"}, )} bare = "I was charged twice for order A-104. Please refund the duplicate." as_list = [m["text"] for m in TICKET["ticket"]["messages"]] shapes = [("string: the message only", bare), ("array : the conversation", as_list), ("object: ticket + order + policy", TICKET)] print(f" {'state shape':6s} input tokens ms") seen = {} for label, state in shapes: response, ms = ask(state, question) seen[label] = response.nouls["eligible"].noul print(f" {label:8s} {'recomputed':>11s} " f"{'score':>6s} {'sum(level*p)':>13s} {'API conf':>9s}") worst = 1.0 for label, text in messages.items(): response, _ = ask(text, {"tone": tone, "urgency": urgency}) t, u = response.choices["tone"], response.scores["urgency"] expected = sum(level * p for level, p in u.probabilities.items()) print(f" {label:10s} {'own call':>10s}") for name, q in FANOUT.items(): single, ms = ask({"postmortem": POSTMORTEM}, {name: q}) seq_ms, seq_tokens = seq_ms + ms, seq_tokens + single.usage.input_tokens a, b = value_of(batched.answers[name]), value_of(single.answers[name]) same = a == b if isinstance(a, str) else abs(a - b) 10s}") if isinstance(a, str) else (lambda v: f"{v:10.3f}") print(f" {name: {seq_ms / batched_ms:.1f}x faster and {seq_tokens / batched_tokens:.1f}x fewer tokens; " f"{agree}/{len(FANOUT)} answers agree, because questions never see each other") return f"{seq_ms / batched_ms:.1f}x faster, {seq_tokens / batched_tokens:.1f}x cheaper, {agree}/{len(FANOUT)} agree" fan_out() Because questions in a request cannot see each other, we can ask everything we might need up front, including questions that only matter on one branch, and read only the relevant answers afterwards. We put ten questions about an incident postmortem, two Choices, two Scores and six Nouls, into one call, then ask each of them again in a call of its own, and compare wall time, input tokens and the answers. The state is sent once instead of ten times, which is where both the latency and the token savings come from, and the agreement column checks the isolation claim directly: a question should receive the same answer whether or not it travels with others. Copy CodeCopiedUse a different Browser INTENT = Choice( instructions="What the user wants the banking assistant to do", criteria={"check_balance": "See a balance or recent transactions", "approve_transfer": "Send or approve a transfer of money", "dispute_charge": "Contest a charge they do not recognise", "close_account": "Close the account permanently", "other": "Anything else, or not clear enough to act on"}, ) STAKES = {"check_balance": 0.50, "dispute_charge": 0.70, "approve_transfer": 0.85, "close_account": 0.90} def route(answer): if answer.choice == "other" or answer.confidence human" bar = STAKES[answer.choice] return f"-> run {answer.choice}" if answer.confidence >= bar else f"-> confirm first (needs {bar:.2f})" @section("5. Confidence-gated routing: the bar rises with the stakes") def gated_routing(): inbox = ["how much is in my checking account", "send 2,000 to my landlord like last month", "i guess maybe move some money around? not sure", "there's a 89.99 charge from a gym i never joined", "shut everything down, i'm done with this bank", "what's the weather like in lisbon"] print(f" {'message':5s} decision") acted = 0 for text in inbox: response, _ = ask(text, {"intent": INTENT}) a = response.choices["intent"] decision = route(a) acted += decision.startswith("-> run") print(f" {text[:50]:15s}" for d in DIMENSIONS) + " (each normalised to 0..1)") for name, row in table.items(): print(f" {name: ".join(f"{n} {sum(w[d] * table[n][d] for d in w):.2f}" for n in ranked)) print("\n Two rankings, four model calls: changing the weights re-ran no inference.") return ", ".join(f"{role}: {who}" for role, who in winners.items()) composite_scoring() Composite scoring keeps the model’s job narrow and the policy explicit. For each candidate we ask four Score questions, each describing concrete situations rather than degrees, normalise every score by its top level, and store the resulting table. The ranking is then plain arithmetic: one weight vector for a senior individual contributor, another for a team lead. Because the judgments are stored separately from the weights, changing what we value instantly re-ranks the candidates and requires no inference. You can trace every position in the ranking back to the dimension that produced it. Copy CodeCopiedUse a different Browser ROOMS = {"living_room": None, "bedroom": None, "kitchen": None, "office": None} def set_lights(room, state): return f"lights in {room} -> {state}" def set_thermostat(room, mode): return f"thermostat in {room} -> {mode}" def play_music(room, genre): return f"playing {genre} in {room}" TOOLS = {"set_lights": (set_lights, "state"), "set_thermostat": (set_thermostat, "mode"), "play_music": (play_music, "genre")} CALL_SPEC = { "tool": Choice(instructions="Which smart-home function the command asks for", criteria={"set_lights": "Turn lights on, off, or dim them", "set_thermostat": "Make a room warmer, cooler, or set eco mode", "play_music": "Play music or audio", "none": "Not a smart-home command this system supports"}), "room": Choice(instructions="Which room the command refers to", criteria=ROOMS), "state": Choice(instructions="If this is a lights command: the requested light state", criteria={"on": None, "off": None, "dim": None}), "mode": Choice(instructions="If this is a thermostat command: the requested mode", criteria={"heat": "Warmer", "cool": "Cooler", "eco": "Energy saving"}), "genre": Choice(instructions="If this is a music command: the requested genre", criteria={"jazz": None, "classical": None, "rock": None, "ambient": None}), } @section("7. Typed function calling, and counting the way Jev can do it") def function_calling(): commands = ["it's freezing in the office, warm it up", "kill the lights in the bedroom", "put on something mellow and jazzy in the kitchen", "order me a pizza"] dispatched = 0 for text in commands: response, ms = ask(text, CALL_SPEC) # every argument asked speculatively, one call c = response.choices tool = c["tool"].choice if tool == "none": print(f" {text!r: no tool (confidence {c['tool'].confidence:.2f})") continue fn, arg = TOOLS[tool] weakest = min(c["tool"].confidence, c["room"].confidence, c[arg].confidence) print(f" {text!r: {tool}(room={c['room'].choice!r}, {arg}={c[arg].choice!r}) " f"weakest judgment {weakest:.2f}, {ms:.0f} ms") print(f" {'': 0.5 for p in probs) print(f" fruits counted: {count} of {len(basket)}") return f"{dispatched}/{len(commands)} commands dispatched from typed answers; counted {count} fruits" function_calling() Function calling becomes a set of closed-set questions: one Choice selects the tool, including an explicit none option for commands we do not support, and one Choice per argument is asked speculatively in the same request. The code reads only the arguments that belong to the selected tool, reports the weakest judgment as the confidence of the whole call, and then executes an ordinary Python function with validated, enumerated values. The second half applies a documented workaround: Jev does not count reliably inside a single question, so we ask one Noul per item in one request and do the sum in code. Copy CodeCopiedUse a different Browser import [truncated for AI cost control]

展開要點與分析

文章情報

工程師中級

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級後設資料。
  • This tutorial provides a complete coding guide to TypeSafe AI's Jev, a System One model designed for non-text, structured judgments. It covers installing the official Python SDK,…

技術影響

可能影響 Agent 架構、工具呼叫、工作流自動化和產品整合。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。