[Submitted on 7 Oct 2026]
Title:Verification and Self-Improvement in Agentic AI: Foundations and Limits
View a PDF of the paper titled Verification and Self-Improvement in Agentic AI: Foundations and Limits, by Chien-Ping Lu
View PDF HTML (experimental)
Abstract:Agentic AI systems can improve by searching longer, receiving additional support, or modifying how they propose and verify outputs. A performance score does not distinguish these mechanisms. We compare these changes through bounded verification with hidden terminal randomness. A stage specifies admissible transcripts, polynomial bounds, an alternating verification protocol, and a terminal checker. Its native reach uses default support; its closure frontier permits all support already admitted by the interface. Under a uniform pointwise probability gap and task-relative soundness, these are well-defined languages. We prove that independent majority amplification preserves both languages, whereas existential acceptance over random tapes can admit incorrect outputs. Exact verification is the zero-randomness case, with placement and completeness results. The randomized-verifier classes satisfy $\Sigma_k^{\mathrm{P}}\subseteq\Sigma_k^{\mathrm{RV}}\subseteq\Sigma_{k+1}^{\mathrm{P}}$; strict enlargement and depth separation require explicit complexity assumptions, while $\mathrm{BPP}=\mathrm{P}$ yields exact companions with the same frontiers. Representation analysis separates invariant acceptance from core-versus-support labels that can change under refactoring. For recursive self-improvement, uniformly bounded self-modification under a common sound interpreter and fixed verification protocol remains within the same verification class. A separate conditional-error budget controls false selection across adaptively chosen candidates. A quota-enforced XOR-synthesis family separates unbounded ratios of search success from changes in the accepted languages; exact and probabilistic audits check the resulting evidence requirements. The framework ties self-improvement claims to obligations on correctness, admissible evidence, verification resources, and selection error.
Comments: 26 pages, 5 figures. Includes proofs and reproducibility artifacts
Subjects:
Artificial Intelligence (cs.AI); Computational Complexity (cs.CC)
Cite as: arXiv:2610.10611 [cs.AI]
(or arXiv:2610.10611v1 [cs.AI] for this version)
https://doi.org/10.48550/arXiv.2610.10611
arXiv-issued DOI via DataCite (pending registration)
Submission history
From: Chien-Ping Lu [view email] [v1] Wed, 7 Oct 2026 06:47:49 UTC (90 KB)
Full-text links:
Access Paper:
View a PDF of the paper titled Verification and Self-Improvement in Agentic AI: Foundations and Limits, by Chien-Ping Lu
View PDF
HTML (experimental)
TeX Source
view license
Ancillary-file links:
Ancillary files (details):
CITATION.cff
LICENSE
MANIFEST.json
README.md
code/exact_audit/audit.json
code/exact_audit/fused_qbf_audit.py
code/exact_audit/patch_agent_audit.py
code/exact_audit/patch_agent_declaration.json
code/exact_audit/test_fused_qbf_audit.py
code/exact_audit/test_patch_agent_audit.py
formalization/StagedAscentFused.lean
formalization/lean-toolchain
integrated/audit.py
integrated/audit_results.json
integrated/figures/amplification.dat
integrated/figures/data_manifest.json
integrated/figures/false_acceptance.dat
integrated/figures/generate_data.py
integrated/figures/xor_success.dat
integrated/test_xor_envelope.py
integrated/xor_envelope.py
integrated/xor_envelope_results.json
reproduce.py
(18 additional files not shown)
Additional Features
Audio Summary
Current browse context:
cs.AI
new | recent | 2026-10
Change to browse by:
cs cs.CC
References & Citations
NASA ADS
Google Scholar
Semantic Scholar
Loading...
Data provided by:
Bibliographic Tools
Bibliographic and Citation Tools
Bibliographic Explorer Toggle
Bibliographic Explorer (What is the Explorer?)
Connected Papers Toggle
Connected Papers (What is Connected Papers?)
Litmaps Toggle
Litmaps (What is Litmaps?)
scite.ai Toggle
scite Smart Citations (What are Smart Citations?)
Code, Data, Media
Code, Data and Media Associated with this Article
alphaXiv Toggle
alphaXiv (What is alphaXiv?)
Links to Code Toggle
CatalyzeX Code Finder for Papers (What is CatalyzeX?)
DagsHub Toggle
DagsHub (What is DagsHub?)
GotitPub Toggle
Gotit.pub (What is GotitPub?)
Huggingface Toggle
Hugging Face (What is Huggingface?)
ScienceCast Toggle
ScienceCast (What is ScienceCast?)
Demos
Demos
Replicate Toggle
Replicate (What is Replicate?)
Spaces Toggle
Hugging Face Spaces (What is Spaces?)
Spaces Toggle
TXYZ.AI (What is TXYZ.AI?)
Related Papers
Recommenders and Search Tools
Link to Influence Flower
Influence Flower (What are Influence Flowers?)
Core recommender toggle
CORE Recommender (What is CORE?)
Author
Venue
Institution
Topic
About arXivLabs
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)