翻訳待ち:Show HN: Tracelint – a linter for AI agent traces, no LLM judge
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。ソース概要:Notifications You must be signed in to change notification settings Fork 0 Star 1 BranchesTags Open more actions menu Latest commit History 40 Commits 40 Commits Folders and files NameName Last commit message Last commi…
AI サービスが一時的に利用できないため、復旧後に翻訳を補完します。
Notifications You must be signed in to change notification settings Fork 0 Star 1 BranchesTags Open more actions menu Latest commit History 40 Commits 40 Commits Folders and files NameName Last commit message Last commit date .github .github examples examples src/tracelint src/tracelint tests tests .gitattributes .gitattributes .gitignore .gitignore CHANGELOG.md CHANGELOG.md CODE_OF_CONDUCT.md CODE_OF_CONDUCT.md CONTRIBUTING.md CONTRIBUTING.md LICENSE LICENSE README.md README.md SECURITY.md SECURITY.md pyproject.toml pyproject.toml Repository files navigation A linter for agent runs — it reads the execution trace of a tool-calling agent (what it actually did) and flags structural bugs deterministically, with the exact evidence and a CI exit code. It runs after the run, on the trace — not on your code — and no second model ever judges it. tracelint reads a tool-calling agent's trace and reports structural defects — schema-violating tool calls, ignored tool errors, hallucinated arguments, loops, and redundant calls — each with the exact trace lines as evidence, and returns a CI exit code. It also ships a fault injector and a per-fault recovery scorecard. Model-as-judge detection of these defects is unreliable (published trace-error benchmarks show low localization accuracy). Many of these defects are structurally decidable and need no judge — that is the entire premise of this tool. No second model ever judges the trace. View the live demo report — the constructed validation suite (one planted instance of every defect, clean controls, and legitimate-but-suspicious cases) plus the robust-vs-buggy recovery scorecard, generated by tracelint demo. Limitations (read first) Deterministic rules catch structural defects, not whether the final answer was correct. Hallucinated-argument, loop, and redundant-call findings are candidates unless structurally proven — legitimate value transforms and intentional retries can trip them; each is shown with its evidence for human review, never asserted as a verdict. High-confidence hallucination detection requires the tool schema to declare field origins (x-value-origin). The recovery scorecard needs labeled task outcomes (success oracles); without them it measures behavioral recovery only ("did not crash"), a weaker claim than correctness. A trace is only as complete as its instrumentation. A rule whose required field is missing is suppressed with a stated reason — tracelint never lints a partial trace as if complete. Quick start The demo runs a keyless validation suite and a recovery scorecard end to end — no API key, no model download: pip install tracelint tracelint demo --html demo.html Lint a trace in CI: tracelint check ./trace.json --tools ./tools.json # exit 2 on a hard_defect Exit codes: 0 clean · 2 a structurally-provable defect (hard_defect) · 3 an input error. Heuristic candidates never fail CI on their own; suppressions are disclosed but are not defects. The rules Rule Finding Tiers R1 schema violation — args fail the tool's JSON Schema hard_defect R2a tool returned an error hard_event (structured signal) / candidate (heuristic) R2b an errored result's value reused by a later side-effecting call hard_defect / candidate R3 hallucinated argument — value not derivable from provenance candidate; hard_defect if the field is annotated provided R4 loop — N identical no-progress calls (polls/retries excluded) candidate R5 redundant call — identical call + identical result, no mutation between candidate R6 malformed arguments — the emitted tool-call arguments are not valid JSON hard_defect R7 unknown tool — a call to a tool absent from the declared toolset (possible hallucinated tool) candidate hard_event and hard_defect are orthogonal to the finding kind: a tool-error event is a hard_event from a structured status field but a candidate from an exception-like string in free-form content. Input format A trace is a JSON object (.json, or .jsonl for many): { "run_id": "run-1", "steps": [ {"type": "message", "role": "user", "content": "cancel order 4521 if it hasn't shipped"}, {"type": "tool_call", "call_id": "c1", "name": "get_order_status", "args": {"order_id": "4521"}}, {"type": "tool_result", "call_id": "c1", "content": {"status": "processing"}, "status": "ok"}, {"type": "tool_call", "call_id": "c2", "name": "cancel_order", "args": {"order_id": "4521", "reason": "not_shipped"}} ], "final": "Order 4521 has been cancelled." } tools.json supplies the ground truth the rules check against: { "tools": { "cancel_order": { "schema": {"type": "object", "properties": {"order_id": {"type": "string"}}, "required": ["order_id"]}, "metadata": {"side_effecting": true} } } } A tool can also declare what failure looks like in its result, so a domain failure returned as a transport success (HTTP 200 carrying {"status": "declined"}) is caught structurally instead of slipping through: { "tools": { "charge_card": { "metadata": { "side_effecting": true, "failure_when": {"pointer": "/status", "in": ["declined", "failed"]} } } } } failure_when is a JSON Pointer into the result plus a match (in / equals / exists); a match is a structured error for R2 (feeding R2a and, on reuse into a side-effecting call, R2b). A side-effecting tool with no failure_when and an unclassifiable result is suppressed with a reason — never counted as a clean pass. The rules run against one canonical trace schema; a thin adapter translates each source's format into it, so the rules never change. Built in: from_openai_messages (OpenAI chat message lists), from_langfuse_trace (a Langfuse trace's observations), and from_otel_spans (OpenTelemetry / OpenInference — the universal standard, so it reaches Arize Phoenix, OpenLLMetry, Langfuse-via-OTel, and datasets like TRAIL, not just one vendor). See examples/langfuse_cookbook.py to lint the traces you already collect in Langfuse and write findings back as scores. On real traces: the adapters are validated against live data, not just the spec — from_langfuse_trace on real Langfuse v4 runs, and from_otel_spans on real TRAIL benchmark traces, where tracelint deterministically localized real tool errors, a malformed tool call, and excessive-retry loops with no model in the loop. Real exports vary, so a new source may need a small adapter tweak — and when a field a rule needs is absent, that rule suppresses (says so) rather than guessing, so an unhandled quirk degrades safely instead of producing a wrong result. More adapters are future work. Lint the traces you already collect check reads native tracelint JSON by default, but --format points it straight at the traces your stack already emits — no manual schema conversion: tracelint check spans.json --format openinference # OTel/OpenInference: Phoenix, OTLP, TRAIL tracelint check messages.json --format openai # an OpenAI chat message list tracelint check trace.json --format langfuse # a Langfuse trace export Most rules need no tool schemas, so this works keyless; add --tools tools.json to light up the schema-dependent rules (R1, and R3's high-confidence tier). A multi-trace input (a .jsonl file, a JSON array, or an OTLP export carrying several trace_ids) fans out to one report each. From the library, the same one-liner: from tracelint import lint_otel_trace report = lint_otel_trace(spans) # spans: your OpenInference span export (a list of dicts) print(report.exit_code) # 0 or 2 See examples/lint_openinference_phoenix.py for an offline, keyless end-to-end run (Phoenix-shaped spans → findings, with and without a tool registry). Straight from a running Arize Phoenix instance: import phoenix as px from tracelint import lint_otel_trace spans = px.Client().get_spans_dataframe().to_dict("records") print(lint_otel_trace(spans).exit_code) Both Phoenix shapes are handled: the span-export JSON (top-level span_kind) and the get_spans_dataframe() records (attributes as attributes.* columns). Recovery scorecard Measure how an agent behaves under injected faults, scored against deterministic success oracles: tracelint scorecard --demo --faults timeout,error,rate_limit --runs 5 The baseline must satisfy the oracle first (else recovery is not measured). Each fault type reports a correctness-recovery rate with a Wilson confidence interval; with no oracle it falls back to behavioral recovery, labeled as weaker. Library from tracelint import lint_trace, default_rules, Trace, ToolRegistry trace = Trace.load("trace.json") registry = ToolRegistry.load("tools.json") report = lint_trace(trace, default_rules(), registry) print(report.exit_code) # 0 or 2 for f in report.active_findings: print(f.rule, f.tier.value, f.summary) Development python -m pytest ruff check src tests The core is dependency-light (jsonschema + stdlib) and the whole test suite is deterministic and offline. A real OpenAI trace-generating agent lives behind the opt-in [real-agent] extra and is never part of the linter. Python 3.10–3.12. Topics Resources Readme MIT license Code of conduct Code of conduct Contributing Contributing Security policy Security policy Activity Stars 1 star Watchers 0 watching Forks 0 forks Report repository