跳到主要内容
AI News HubLIVE
站内改写6 分钟阅读

待翻译:A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model

文章摘要

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:This tutorial provides a complete coding guide to TypeSafe AI's Jev, a System One model designed for non-text, structured judgments. It covers installing the official Python SDK, using primitive question types (Choice, Score, Noul), implementing speculative fan-out, confidence-gated routing, and building async production workflows The post A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model appeared first on MarkTechPost.

来源MarkTechPost作者: Asif Razzaq
待翻译:A Coding Guide to TypeSafe AI Jev: Typed Decisions, Calibrated Confidence, and Speculative Fan-Out with a System One Model
报告错误

纠错通道尚未开通,可先复制下方文章信息留存。

查看更正说明
直接读正文

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

In this tutorial, we work with Jev, TypeSafe AI’s first System One model, which does not generate text at all: we send it a piece of program state and a set of typed questions, and it returns choices, scores, and yes/no probabilities that our code can branch on directly. We install the official Python SDK, make a first call that uses all three question primitives at once, and look at how the shape of the state changes what the model can know. We then recompute the published confidence statistic from the returned probabilities, measure what batching ten questions into one call buys over ten separate calls, and build the patterns the API is designed for: confidence-gated routing, composite scoring with the weights kept in code, typed function calling, and counting done the way the model can actually do it. We close with the production shape: Pydantic response models, an async client fanned out with asyncio, retry policies, typed errors, and a running ledger that prices the whole notebook. Copy CodeCopiedUse a different Browser import os import sys import json import time import asyncio import traceback import subprocess from getpass import getpass RESULTS = {} LEDGER = {"calls": 0, "input_tokens": 0, "output_tokens": 0} USD_PER_MILLION_INPUT_TOKENS = 0.042 # Jev list price; output tokens are free def banner(title): print("\n" + "=" * 78) print(title) print("=" * 78) def section(name): def wrap(fn): def run(*a, kw): banner(name) try: out = fn(*a, kw) RESULTS[name] = out if isinstance(out, str) else "ok" return out except Exception as e: RESULTS[name] = f"SKIPPED / FAILED -> {type(e).name}: {e}" print(f"\n[!] {name} did not complete: {type(e).name}: {e}") traceback.print_exc(limit=3) return None return run return wrap banner("0. Install the SDK, load the API key, list the models") subprocess.run([sys.executable, "-m", "pip", "install", "-q", "typesafe-sdk==0.7.0"], check=True) import typesafe_sdk from typesafe_sdk import Choice, Noul, Score, TypeSafeClient def load_api_key(): key = os.environ.get("TYPESAFE_API_KEY", "").strip() if not key: try: from google.colab import userdata # Colab: key stored under the Secrets tab key = (userdata.get("TYPESAFE_API_KEY") or "").strip() except Exception: key = "" return key or getpass("TypeSafe API key (console.typesafe.ai/keys): ").strip() os.environ["TYPESAFE_API_KEY"] = load_api_key() client = TypeSafeClient() # reads TYPESAFE_API_KEY, defaults to jev-latest print(f" typesafe-sdk {typesafe_sdk.version} | Python {sys.version.split()[0]}") print(" models available to this key:") for m in client.models.list().models: print(f" {m.name: {dept.choice!r} confidence {dept.confidence:.3f}") print(f" probabilities {({k: round(v, 3) for k, v in dept.probabilities.items()})}") fr = response.scores["frustration"] print(f" frustration -> score {fr.score:.3f} on 0..{len(fr.legend) - 1} confidence {fr.confidence:.3f}") for level, text in fr.legend.items(): print(f" {level}: p={fr.probabilities[level]:.3f} {text}") print(f" refund_requested -> noul {response.nouls['refund_requested'].noul:.3f}") print(f" policy_supports -> noul {response.nouls['policy_supports'].noul:.3f}") print(f"\n answered by {response.model} in {ms:.0f} ms " f"input tokens {response.usage.input_tokens}, output tokens {response.usage.output_tokens}") return f"{dept.choice}, frustration {fr.score:.2f}, refund {response.nouls['refund_requested'].noul:.2f}" three_primitives() A System One request has two parts: state, which is any text, JSON object or array describing the situation, and a dictionary of named questions. Choice selects one label from the criteria we define and returns a probability for every label; Score places the state on an ordered rubric and returns the probability-weighted level, so it can land between two levels; Noul returns a single probability that a statement is true. The question names are ours and never reach the model, which is why the instructions carry the full meaning and can point at nested fields with backticked paths. All four questions are evaluated in one request, in parallel and in isolation from one another, and the response reports the pinned model version that answered and the tokens it billed. Copy CodeCopiedUse a different Browser @section("2. State is program state: the same question over a string and over named fields") def state_shapes(): question = {"eligible": Noul( instructions="The customer is eligible for a refund under the company's written policy", criteria={"true": "A policy is present and it covers the customer's situation", "false": "No policy is given, or the policy does not cover the situation"}, )} bare = "I was charged twice for order A-104. Please refund the duplicate." as_list = [m["text"] for m in TICKET["ticket"]["messages"]] shapes = [("string: the message only", bare), ("array : the conversation", as_list), ("object: ticket + order + policy", TICKET)] print(f" {'state shape':6s} input tokens ms") seen = {} for label, state in shapes: response, ms = ask(state, question) seen[label] = response.nouls["eligible"].noul print(f" {label:8s} {'recomputed':>11s} " f"{'score':>6s} {'sum(level*p)':>13s} {'API conf':>9s}") worst = 1.0 for label, text in messages.items(): response, _ = ask(text, {"tone": tone, "urgency": urgency}) t, u = response.choices["tone"], response.scores["urgency"] expected = sum(level * p for level, p in u.probabilities.items()) print(f" {label:10s} {'own call':>10s}") for name, q in FANOUT.items(): single, ms = ask({"postmortem": POSTMORTEM}, {name: q}) seq_ms, seq_tokens = seq_ms + ms, seq_tokens + single.usage.input_tokens a, b = value_of(batched.answers[name]), value_of(single.answers[name]) same = a == b if isinstance(a, str) else abs(a - b) 10s}") if isinstance(a, str) else (lambda v: f"{v:10.3f}") print(f" {name: {seq_ms / batched_ms:.1f}x faster and {seq_tokens / batched_tokens:.1f}x fewer tokens; " f"{agree}/{len(FANOUT)} answers agree, because questions never see each other") return f"{seq_ms / batched_ms:.1f}x faster, {seq_tokens / batched_tokens:.1f}x cheaper, {agree}/{len(FANOUT)} agree" fan_out() Because questions in a request cannot see each other, we can ask everything we might need up front, including questions that only matter on one branch, and read only the relevant answers afterwards. We put ten questions about an incident postmortem, two Choices, two Scores and six Nouls, into one call, then ask each of them again in a call of its own, and compare wall time, input tokens and the answers. The state is sent once instead of ten times, which is where both the latency and the token savings come from, and the agreement column checks the isolation claim directly: a question should receive the same answer whether or not it travels with others. Copy CodeCopiedUse a different Browser INTENT = Choice( instructions="What the user wants the banking assistant to do", criteria={"check_balance": "See a balance or recent transactions", "approve_transfer": "Send or approve a transfer of money", "dispute_charge": "Contest a charge they do not recognise", "close_account": "Close the account permanently", "other": "Anything else, or not clear enough to act on"}, ) STAKES = {"check_balance": 0.50, "dispute_charge": 0.70, "approve_transfer": 0.85, "close_account": 0.90} def route(answer): if answer.choice == "other" or answer.confidence human" bar = STAKES[answer.choice] return f"-> run {answer.choice}" if answer.confidence >= bar else f"-> confirm first (needs {bar:.2f})" @section("5. Confidence-gated routing: the bar rises with the stakes") def gated_routing(): inbox = ["how much is in my checking account", "send 2,000 to my landlord like last month", "i guess maybe move some money around? not sure", "there's a 89.99 charge from a gym i never joined", "shut everything down, i'm done with this bank", "what's the weather like in lisbon"] print(f" {'message':5s} decision") acted = 0 for text in inbox: response, _ = ask(text, {"intent": INTENT}) a = response.choices["intent"] decision = route(a) acted += decision.startswith("-> run") print(f" {text[:50]:15s}" for d in DIMENSIONS) + " (each normalised to 0..1)") for name, row in table.items(): print(f" {name: ".join(f"{n} {sum(w[d] * table[n][d] for d in w):.2f}" for n in ranked)) print("\n Two rankings, four model calls: changing the weights re-ran no inference.") return ", ".join(f"{role}: {who}" for role, who in winners.items()) composite_scoring() Composite scoring keeps the model’s job narrow and the policy explicit. For each candidate we ask four Score questions, each describing concrete situations rather than degrees, normalise every score by its top level, and store the resulting table. The ranking is then plain arithmetic: one weight vector for a senior individual contributor, another for a team lead. Because the judgments are stored separately from the weights, changing what we value instantly re-ranks the candidates and requires no inference. You can trace every position in the ranking back to the dimension that produced it. Copy CodeCopiedUse a different Browser ROOMS = {"living_room": None, "bedroom": None, "kitchen": None, "office": None} def set_lights(room, state): return f"lights in {room} -> {state}" def set_thermostat(room, mode): return f"thermostat in {room} -> {mode}" def play_music(room, genre): return f"playing {genre} in {room}" TOOLS = {"set_lights": (set_lights, "state"), "set_thermostat": (set_thermostat, "mode"), "play_music": (play_music, "genre")} CALL_SPEC = { "tool": Choice(instructions="Which smart-home function the command asks for", criteria={"set_lights": "Turn lights on, off, or dim them", "set_thermostat": "Make a room warmer, cooler, or set eco mode", "play_music": "Play music or audio", "none": "Not a smart-home command this system supports"}), "room": Choice(instructions="Which room the command refers to", criteria=ROOMS), "state": Choice(instructions="If this is a lights command: the requested light state", criteria={"on": None, "off": None, "dim": None}), "mode": Choice(instructions="If this is a thermostat command: the requested mode", criteria={"heat": "Warmer", "cool": "Cooler", "eco": "Energy saving"}), "genre": Choice(instructions="If this is a music command: the requested genre", criteria={"jazz": None, "classical": None, "rock": None, "ambient": None}), } @section("7. Typed function calling, and counting the way Jev can do it") def function_calling(): commands = ["it's freezing in the office, warm it up", "kill the lights in the bedroom", "put on something mellow and jazzy in the kitchen", "order me a pizza"] dispatched = 0 for text in commands: response, ms = ask(text, CALL_SPEC) # every argument asked speculatively, one call c = response.choices tool = c["tool"].choice if tool == "none": print(f" {text!r: no tool (confidence {c['tool'].confidence:.2f})") continue fn, arg = TOOLS[tool] weakest = min(c["tool"].confidence, c["room"].confidence, c[arg].confidence) print(f" {text!r: {tool}(room={c['room'].choice!r}, {arg}={c[arg].choice!r}) " f"weakest judgment {weakest:.2f}, {ms:.0f} ms") print(f" {'': 0.5 for p in probs) print(f" fruits counted: {count} of {len(basket)}") return f"{dispatched}/{len(commands)} commands dispatched from typed answers; counted {count} fruits" function_calling() Function calling becomes a set of closed-set questions: one Choice selects the tool, including an explicit none option for commands we do not support, and one Choice per argument is asked speculatively in the same request. The code reads only the arguments that belong to the selected tool, reports the weakest judgment as the confidence of the whole call, and then executes an ordinary Python function with validated, enumerated values. The second half applies a documented workaround: Jev does not count reliably inside a single question, so we ask one Noul per item in one request and do the sum in code. Copy CodeCopiedUse a different Browser import [truncated for AI cost control]

展开要点与分析

文章情报

工程师中级

要点

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • This tutorial provides a complete coding guide to TypeSafe AI's Jev, a System One model designed for non-text, structured judgments. It covers installing the official Python SDK,…

要点与分析由自动化流程生成,可能有误,请结合原始来源核实。