跳到主要内容
AI News HubLIVE
来源内容 · 翻译待补全2 分钟阅读

待翻译:Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference

文章摘要

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Architect Financial Technologies has launched Liquid Inference, an LLM router that runs a live auction for every request. Liquid Inference is an LLM inference marketplace from Architect where providers bid to serve each prompt. The buyer pays the lowest offer that meets its rules. For developers, it is quite simple message: swap a base URL, […] The post Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference appeared first on MarkTechPost.

来源MarkTechPost作者: Michal Sutter
待翻译:Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference
报告错误

纠错通道尚未开通,可先复制下方文章信息留存。

查看更正说明
直接读正文

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

Architect Financial Technologies has launched Liquid Inference, an LLM router that runs a live auction for every request. Liquid Inference is an LLM inference marketplace from Architect where providers bid to serve each prompt. The buyer pays the lowest offer that meets its rules. For developers, it is quite simple message: swap a base URL, keep your code, and let providers compete on price. What is Liquid Inference? Liquid Inference is an exchange-style router for LLM inference. According to Architect, providers post offers to serve specific models. Each request is auctioned across every provider quoting the named model. The lowest-priced offer that meets the buyer’s rules wins. The product comes from a trading firm, not an AI lab. Architect runs the AX perpetual futures exchange. In May 2026 it acquired a US Designated Contract Market to list GPU compute futures, pending regulatory review. The team used its experience building financial exchanges to create two-sided price discovery for inference. How does the inference auction work? The flow has 4 steps: Request: A client sends a standard OpenAI or Anthropic API call. Rules: The buyer’s constraints filter eligible offers. Auction: Providers quoting that model compete. The lowest qualifying offer wins. Receipt: The max price is locked before generation. Billing covers metered usage only. Buyers can set per-job cost caps, time to first token limits and minimum throughput. They can also require approved regions, zero data retention and provider or model allow lists. An Auto mode can pick the model for a given unit of work. Harrison’s LinkedIn post adds routing rule presets and full multi-modal support. Account holders can view live order books, per-provider and per-model quotes, and cleared transactions. That level of market data is unusual for an LLM API. What do buyers get? Drop-in compatibility with agentic coding tools. The post lists Claude Code, Codex, OpenCode, Cursor, Pi and Cline. Free email signup. The first 500 users get $20 of free inference, per Harrison. A referral program: 20% of referred fees as free inference, plus 10% on second-level referrals. What do inference providers get? Providers onboard through the Liquid Inference app. Harrison says new providers are verified “in minutes, not weeks.” All prompts use the OpenAI API standard. A REST and WebSocket API registers models and quotes. Providers can update quotes based on their own costs. That lets them sell spare GPU capacity only when they want. Payouts run through Stripe, with itemized records of every job. How does Liquid Inference compare with OpenRouter and Hugging Face? FeatureLiquid InferenceOpenRouterHugging Face Inference Providers Routing modelPer-request auction across quoting providers Price-weighted load balancing, inverse square of price Fastest provider by default; :cheapest suffix optional Model / provider count“Hundreds” of models; providers not disclosed500+ models, 80+ providers 18 listed partners API compatibilityOpenAI and Anthropic-compatibleOpenAI-compatibleOpenAI-compatible, chat only Price capMax price locked before first tokenmax_price parameter Not disclosed Data controlsZDR, regions, allow listszdr, data_collection, only/ignore Provider preference order Platform feeNot disclosed5.5% card credit fee, $0.80 minimum No markup Free credits$20 for first 500 users Not disclosed$0.10/month free, $2.00 PRO Public market dataLive order books and cleared tradesNot disclosedNot disclosed Key Takeaways Liquid Inference auctions every LLM request across competing providers. Max price is locked before the first token, with a per-job receipt. Works with OpenAI and Anthropic clients, including Claude Code and Cursor. First 500 users get $20 free; referrals earn 20% of fees. Fees, provider list and latency data are not yet public. Check out the Technical details. All credit goes to the researcher of this project. Also, feel free to follow us on Twitter and don’t forget to join our 150k+ML SubReddit and Subscribe to our Newsletter. Wait! are you on telegram? now you can join us on telegram as well. The post Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference appeared first on MarkTechPost.

展开要点与分析

文章情报

工程师进阶

要点

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Architect Financial Technologies has launched Liquid Inference, an LLM router that runs a live auction for every request. Liquid Inference is an LLM inference marketplace from Arc…

要点与分析由自动化流程生成,可能有误,请结合原始来源核实。