AI News HubLIVE
站內改寫2 分鐘閱讀

待翻譯:GLM 5.3 Flash faster and cheaper

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:GLM 5.3 Flash | Model APIs | RunInfra RunInfraby RightNow © 2026 RunInfra. All rights reserved. Join the communitySystem status Backed by Combinator AICPA Type II SOC 2 Ask AI about RunInfra Part of RightNow RunInfraby…

來源Hacker News AI作者: OsamaJaber

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

GLM 5.3 Flash | Model APIs | RunInfra RunInfraby RightNow © 2026 RunInfra. All rights reserved. Join the communitySystem status Backed by Combinator AICPA Type II SOC 2 Ask AI about RunInfra Part of RightNow RunInfraby RightNow © 2026 RunInfra. All rights reserved. Join the communitySystem status Backed by Combinator AICPA Type II SOC 2 Ask AI about RunInfra Part of RightNow GLM 5.3 Flash zai-org/GLM-5.3-Flash Get API keyView docs GLM 5.3 Flash is an LLM listed in RunInfra Model APIs. RunInfra serves it as zai-org/GLM-5.3-Flash at $0.10 per 1M input tokens and $0.40 per 1M output tokens. Its context window is 1,048,576 tokens. The API provides OpenAI-compatible chat completions. Pricing USD, pay per token per 1M input tokens$0.10 per 1M cached input tokens$0.01 per 1M output tokens$0.40 Measured performance Output speed 254.1output tokens per second, model only Time to first token 703milliseconds to first reasoning token Cache hit rate 81%of input tokens served from the cache Access Confirm how your client reaches this model. ProviderZ.ai API compatibilityOpenAI-compatible chat completions Accepted inputText and images AvailabilityAvailable Data retentionYour prompts are never stored and never used for training. Code examples Set RUNINFRA_GATEWAY_KEY to your workspace API key before using an example. Trust and provenance Verify the company and operating credentials behind this API. RunInfra is a sub-product of RightNow Research Lab. SOC 2 Type IIAudited access, logging, and incident response. RunInfraby RightNow © 2026 RunInfra. All rights reserved. Join the communitySystem status Capacity Check the limits your workload must fit. Context window1,048,576 tokens Maximum request size3.5 MB per request Maximum generated output1,048,576 tokens Capabilities See which request modes the API supports. Tool callingSupported JSON modeSupported StreamingSupported PrecisionFP8, vendor-native release View Full Spec Gateway compatibilityOpenAI-compatible chat completions for compatible clients and gateways OpenRouter Provider MonitorProvider metadata published as ready for discovery. Not proof of a live OpenRouter listing or callability. Prefix cachingAutomatic prefix caching runs on every replica, and requests from the same session or conversation are routed back to the replica that holds their cached prefix. A stable session hint or prompt_cache_key provides explicit grouping; otherwise RunInfra derives it from the start of the conversation. Hits remain best effort until eviction or a serve restart, never guaranteed. Cached input is billed at the cached input rate. Responses report cached token counts. Cache retentionPrefix cache is held on the serving GPU and evicted under memory pressure. Retention is best effort. Upstream modelView model Y CombinatorBacked by Y Combinator. NVIDIA InceptionMember of NVIDIA Inception. Optimization agent Model APIs Pricing Startups Benchmarks Docs Research News Contact Backed by Combinator AICPA Type II SOC 2 Ask AI about RunInfra Part of RightNow