待翻译:A 10M-token context efficient agentic model
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Pokee-Isaac 28B | Pokee Console Loading page Model at a glance 28Bparameters The smallest model in the comparison panel, by a wide margin. 10Mtoken context Usable end to end, not merely addressable. $0.15 / $1.00per 1M…
AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。
Pokee-Isaac 28B | Pokee Console Loading page Model at a glance 28Bparameters The smallest model in the comparison panel, by a wide margin. 10Mtoken context Usable end to end, not merely addressable. $0.15 / $1.00per 1M in / out Below every baseline that can be bought at this length. Benchmark overview Every benchmark, every model, one table. The full comparison at a glance. Each section below takes one row of this table and shows how the result was reached — where the panel diverges, and where a baseline finishes ahead. API pricing: USD per 1M input / output tokens Pokee-Isaac 28B compared with five baseline models across every benchmark in the technical report, with per-token pricing for each model BenchmarkPokee-Isaac28B (v0)In/Out$0.15 / $1.00 GPT-5.6 LunaAzureIn/Out$0.40 / $1.80>272K contextGemini 3.5 Flash LiteVertex AIIn/Out$0.30 / $2.50 Claude Haiku 4.5BedrockIn/Out$1.00 / $5.00 Nemotron 3 SuperAmazon Bedrock · USIn/Out$0.15 / $0.65 Qwen 3.5 122BOpenRouterIn/Out$0.26 / $2.08 Long context RULER256K / 512K / 1M↑ higher is better96.9 / 96.7 / 95.0🥇 (best in row)95.0 / 91.4 / 0.0*94.5 / 94.6 / 29.4*0.0 / 0.0 / 0.096.3 / 95.7 / 91.8s0.0 / 0.0 / 0.0 RULER2M / 4M / 10M↑ higher is better95.8 / 96.7 / 93.3🥇 (best in row)0.0 / 0.0 / 0.00.0 / 0.0 / 0.00.0 / 0.0 / 0.00.0 / 0.0 / 0.00.0 / 0.0 / 0.0 MRCR v2256K / 512K / 1M, 8 needles↑ higher is better0.607 / 0.743 / 0.500🥇 (best in row)0.208 / 0.173 / 0.0500.474 / 0.473 / 0.2050.000 / 0.000 / 0.0000.145 / 0.161 / 0.0670.000 / 0.000 / 0.000 Agentic capabilities BFCL v4overall↑ higher is better70.94🥇 (best in row)70.6164.8567.5233.1364.88 τ³-bench4-domain average↑ higher is better0.662🥇 (best in row)0.5270.6310.4080.4260.611 Terminal-Bench 2.1text-only subset↑ higher is better65.1%69.8%🥇 (best in row)46.5%34.9%24.4%46.5% MCP-Atlasclaim coverage↑ higher is better74.59%77.90%🥇 (best in row)76.67%56.45%48.95%70.24% Security DTAPattack success rate (ASR)↓ lower is better35.6🥇 (best in row)50.166.337.960.454.0 DTAPbenign success rate (BSR)↑ higher is better82.585.1🥇 (best in row)83.371.363.379.4 🥇 marks the best value in each row. ASR is attack success rate and is lower-is-safer; BSR is benign task success rate; higher is better for every other benchmark. A 0.0 / 0.000 means the model returned nothing usable at that length. * context-overflow error at 1M. s vendor self-reported, not measured by Pokee — excluded from the row comparison. Every other figure was produced on one installation, for Isaac and each baseline alike. Pricing: Pokee and Luna rates are as supplied/official; Gemini and Haiku use standard public API rates; Nemotron uses Amazon Bedrock on-demand pricing for US East / US West; Qwen uses the OpenRouter headline rate, which varies by provider. Haiku, Nemotron, and Qwen cannot be purchased at the context lengths above 262K that this comparison covers. Long context A window that is usable, not merely advertised. RULER holds task difficulty fixed and scales only the context length, so the curve shows how far a model's usable context tracks its nominal one. Isaac is the only model in the panel that returns a score at every length. RULER score by context length for Pokee-Isaac 28B and five baseline models Model256K512K1M2M4M10M Pokee-Isaac 28B96.996.795.095.896.793.3 GPT-5.6 Luna(Azure)95.091.4—0.0, context-overflow error—0.0, no usable score—0.0, no usable score—0.0, no usable score Gemini 3.5 Flash Lite(Vertex AI)94.594.629.4* (context-overflow error)—0.0, no usable score—0.0, no usable score—0.0, no usable score Claude Haiku 4.5(Bedrock)—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score Nemotron 3 Super 120B96.3 (self-reported by vendor)95.67 (self-reported by vendor)91.75 (self-reported by vendor)—0.0, no usable score—0.0, no usable score—0.0, no usable score Qwen 3.5 122B—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score—0.0, no usable score Pokee-IsaacBaselineVendor self-reported—No usable score (0.0) RULER score (%, averaged across task configurations), ten samples per configuration. The 256K and 512K columns average all 13 configurations; common-words extraction is unavailable from 1M onward, so those columns average the remaining 12. * marks a context-overflow error at that length. Nemotron's 256K–1M figures are self-reported by NVIDIA, not measured by Pokee, and are excluded from the comparison. MRCR v2 — interference, not just depth RULER measures how deep a model can reach; MRCR measures whether it can tell several buried targets apart. GPT-5.6 Luna scores 95.0 on RULER at 256K but 0.050 here at 1M — multi-needle disambiguation splits the panel far more sharply than single-target recall does. Leads at every length 256K +0.133 vs. next Pokee-Isaac0.607 Gemini 3.5 Flash Lite0.474 GPT-5.6 Luna0.208 Nemotron 3 Super0.145 Claude Haiku 4.50.000 Qwen 3.5 122B0.000 512K +0.270 vs. next Pokee-Isaac0.743 Gemini 3.5 Flash Lite0.473 GPT-5.6 Luna0.173 Nemotron 3 Super0.161 Claude Haiku 4.50.000 Qwen 3.5 122B0.000 1M +0.295 vs. next Pokee-Isaac0.500 Gemini 3.5 Flash Lite0.205 Nemotron 3 Super0.067 GPT-5.6 Luna0.050 Claude Haiku 4.50.000 Qwen 3.5 122B0.000 MRCR v2 (Multi-Round Co-reference Resolution), 8-needle split, 0–1 scale, all three panels on the same axis. Multiple needles are distributed through a long conversation and the model must retrieve and disambiguate a specified one, so partial recall and cross-needle interference are both penalized. Isaac leads at every measured length and the margin over the closest baseline widens as context grows. A score of 0.000 means the model returned nothing usable at that length. Agentic capability Four benchmarks, four different facets of agency. Each probes something the others cannot mask: deterministic function calling, sustained multi-turn coherence, execution in a real shell, and discovery across live tool servers. Isaac leads two, places second on one, and third on one. BFCL v4 Function calling, scored by AST and state-transition matching rather than an LLM judge. 1st of 6 Pokee-Isaac70.94 GPT-5.6 Luna70.61 Claude Haiku 4.567.52 Qwen 3.5 122B64.88 Gemini 3.5 Flash Lite64.85 Nemotron 3 Super33.13 Overall score across all 5,217 examples. The 0.33-point margin over GPT-5.6 Luna is parity, not a decisive lead — the substantive claim is matching the strongest cloud baseline while staying deployable in boundary. τ³-bench Multi-turn customer-service tasks against a simulated user whose requirements evolve. 1st of 6 Pokee-Isaac0.6620.820 / 3-dom Gemini 3.5 Flash Lite0.6310.774 / 3-dom Qwen 3.5 122B0.6110.767 / 3-dom GPT-5.6 Luna0.5270.641 / 3-dom Nemotron 3 Super0.4260.557 / 3-dom Claude Haiku 4.50.4080.524 / 3-dom Overall average across all four domains — retail, airline, telecom, and banking. The three-domain average excluding banking is shown alongside; banking is where the whole panel struggles, with the best score in that column at 0.203 and Isaac at 0.186. Terminal-Bench 2.1 Hard, realistic tasks driven through a real shell, verified by the task's own test suite. 2nd of 6 GPT-5.6 Luna69.8%60/86 Pokee-Isaac65.1%56/86 Gemini 3.5 Flash Lite46.5%40/86 Qwen 3.5 122B46.5%40/86 Claude Haiku 4.534.9%30/86 Nemotron 3 Super24.4%21/86 Text-only subset, 86 of 89 tasks, 100-turn cap. This is the one benchmark in the report where a cloud baseline finishes ahead of Isaac — a gap of four tasks. MCP-Atlas Discovering and composing tools across 36 live MCP servers, with neither server nor tool named. 3rd of 6 GPT-5.6 Luna77.9%6.91 turns Gemini 3.5 Flash Lite76.7%14.99 turns Pokee-Isaac74.6%9.10 turns Qwen 3.5 122B70.2%11.47 turns Claude Haiku 4.556.5%7.40 turns Nemotron 3 Super49.0%11.18 turns Mean claim coverage over 500 public tasks. Isaac lands within 2.1 points of Gemini while taking roughly 60% of its trajectory length. Security Safest of the panel on both attack axes. DTAP places an agent in simulated environments and measures whether injected attacks succeed. Against the five baselines run on the same installation with the judge held fixed, Isaac is the safest of six on both axes while placing third of six on capability. DTAP red-teaming results, macro-averaged over the 12 Linux domains, same judge for every row ModelDirect ASRlower is saferIndirect ASRlower is saferCombined ASRlower is saferBSRhigher is better Pokee-Isaac 28B36.0best (best in column)35.2best (best in column)35.6best (best in column)82.5 GPT-5.6 Luna54.446.150.185.1best (best in column) Gemini 3.5 Flash Lite84.149.566.383.3 Claude Haiku 4.538.237.137.971.3 Nemotron 3 Super 120B80.342.060.463.3 Qwen 3.5 122B60.847.854.079.4 DTAP (DecodingTrust-Agent Platform) places an agent in simulated environments and measures whether injected attacks succeed. Direct ASR is the attack success rate when the harmful request is in the user prompt; Indirect ASR when it arrives through tool output or the environment. BSR is benign task success — utility. Macro-average over 12 Linux domains and 6,195 judged tasks, same judge for every row. Isaac is the safest of the six on both attack axes, and its Direct and Indirect rates differ by 0.8 points — the tightest balance in the set. Two limits are stated plainly in the report: the indirect figures are a guards-inactive measurement, because the harness matches native tool names while DTAP's attacks arrive over MCP; and explicit refusals fired on only 1.5% of malicious tasks, so most of the current defense is incidental rather than declined. Against the broader published leaderboard of sixteen systems, Isaac places #5 on Direct ASR, #6 on Indirect, and #9 on capability — a middle-of-the-field result, and the report states it as one. Efficiency Prefill throughput rises with context length. Measured under the same RULER workload whose accuracy is reported above, on a single B200-class GPU — so the serving profile is directly comparable to the capability curve rather than benchmarked on a friendlier task. 137,200peak prefill tok/s at 10M Top of the measured sweep on a single NVIDIA B200; 42,400 at 1M. 72.9 stime to first token For a full 10M-token prompt on that B200. ~335decode tok/s, flat Unchanged from 1M to 10M context. Serving profile under the RULER workload on a single B200-class GPU ContextConcurrencyTTFTPrefill (tok/s)Decode (tok/s) 1M123.6 s42,400335 1M449.3 s81,200322 10M172.9 s137,200337 Pricing Cheaper on both meters, across a window an order of magnitude larger. Published rates per million tokens. Three models in the panel carry a rate but cannot be bought at the context lengths this evaluation covers — no commercially available endpoint serves them beyond 262K. List pricing in USD per million tokens at long-context lengths ModelMax contextInput $/MOutput $/M Pokee-Isaac 28B10M$0.15$1.00 GPT-5.6 LunaAzure1.05M$0.40$1.80>272K context Gemini 3.5 Flash LiteVertex AI1M$0.30$2.50 Claude Haiku 4.5Bedrock200Knot sold above 262K$1.00$5.00 Nemotron 3 Super 120BAmazon Bedrock · US262Knot sold above 262K$0.15$0.65 Qwen 3.5 122BOpenRouter262Knot sold above 262K$0.26$2.08 Add credits Retrieved from each provider's public pricing page on 3 August 2026; long-context rates are quoted where a provider meters them separately. Nemotron uses Amazon Bedrock on-demand pricing for US East / US West; Qwen uses the OpenRouter headline rate, which varies by provider. Pokee-Isaac rates are provisional and subject to confirmation at launch. For in-boundary deployments the per-token comparison understates the difference — cost becomes a fixed function of the hardware provisioned rather than a variable function of tokens consumed. Deployment From a datacenter GPU down to a single consumer card. Portability is a first-class property: Isaac is adapted to run on heterogeneous accelerator hardware, [truncated for AI cost control]