待翻译:Show HN: InferCrane – One stable endpoint for self-hosted AI inference
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Open-source production inference Give InferCrane a model. Get a production endpoint. Deploy a model or connect what already runs. InferCrane operates it behind one stable endpoint. Join the Cloud waitlistView on GitHub…
AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。
Open-source production inference Give InferCrane a model. Get a production endpoint. Deploy a model or connect what already runs. InferCrane operates it behind one stable endpoint. Join the Cloud waitlistView on GitHub → Apache-2.0 · runs in your infrastructure · InferCrane Cloud in private preview Model in. Stable endpoint out. You giveYour model or endpointopen-weight · custom · existing InferCraneplans and operates Stable endpointcoder-productionOpenAI-compatible · healthy Serving planreplaceableruntime · GPU · provider Candidatemeasured firstbenchmark · replay · quality Start with a model and the outcome you need.intent received Three ways to begin Deploy new. Connect existing. Govern model APIs. Start with what your team has today. Keep one application interface as the infrastructure behind it changes. 01Start from a model Deploy new inference Choose a model and objective. InferCrane plans and runs the deployment. Model → plan → durable deployment → endpoint 02Keep what works Connect existing inference Observe vLLM, SGLang, LiteLLM, or another endpoint before changing traffic. Observe → route → manage · no forced migration 03Use external capacity Govern a model API Keep provider credentials server-side and enforce privacy and spending limits. Stable identity · controlled spend · portable future Model or existing endpoint→model="support-production"→Operated production inferenceYour application keeps the same model identity. Cost and performance without guesswork Know when owning inference is actually worth it. Compare model APIs and self-hosted plans using measured latency, throughput, reliability, and sourced cost. Start nowConnect Model API Reach first traffic quickly through an approved OpenAI-compatible provider. Provider usage · latency · errors · budget → Prove the alternativeMeasure Self-hosted candidate Benchmark an exact artifact, runtime, GPU, provider, and workload shape. TTFT · TPOT · throughput · errors · sourced cost → Choose by policyDecide Production serving plan Promote owned capacity, retain governed fallback, or keep the API binding. Release decision · hard budget · audit trail Every recommendation shows its evidenceMeasuredModeledNot yet measuredNo invented savings or prices. Built for production workloads One inference system for the applications you are building. Start with an API, open model, private fleet, agent, or retrieval workload. Keep the same operating model as the application grows. API-first AI applications Start with an API. Own inference when it makes sense. Keep one application model identity across model APIs, gateways, and customer-operated capacity. Move workloads only when measured cost, latency, or privacy evidence justifies the change. ✓ One model alias · hard API budgets · measured migrationSee the workflow → $client.responses.create($ model="support-production",$ input="Summarize this customer conversation"$) One product across the inference lifecycle Your application stays simple. The infrastructure can evolve. InferCrane owns deployment, policy, measurement, and release decisions. Runtimes, gateways, and compute remain replaceable. Explore the architecture → InferCrane operated boundary Healthy Your applicationmodel="coder-production"unchanged Stable endpointcoder-productionPOST /v1/responses InferCrane operates what changesdesired state → measured state 01 / ModelImmutable artifactweights · tokenizer · provenance selected artifactcached 02 / RuntimeQualified serving planengine · precision · batching vLLM+L40S 03 / CapacityDemand-aware replicasadmission · scale · drain RouteObserveOptimizeReleaseRecover Replaceable execution AWSGCPKubernetesExisting stack Evidence carried with every change latencyqualitycosthealth Works with your inference stack Use proven infrastructure without stitching it together yourself. InferCrane coordinates runtimes, gateways, compute, and measurement tools behind one production endpoint. Serve vLLMSGSGLangOCICustom OCI Connect LLLiteLLMOpenRouterAPIOpenAI-compatible APIs Compute AWSAWSGoogle CloudKubernetesRPRunPod Measure OpenTelemetryAIAIPerf Before you try it Common questions, direct answers. Can I connect inference I already run? Yes. Start in observe-only mode with vLLM, SGLang, LiteLLM, or another OpenAI-compatible endpoint. Add traffic or lifecycle ownership only when your team is ready. Does InferCrane replace my serving engine or cloud? No. InferCrane operates around proven runtimes, gateways, and infrastructure. Those execution layers remain replaceable while your application keeps one model identity. Where does the compute run? Self-hosted workloads run in infrastructure you control or behind external endpoints you connect. Provider credentials stay on the control-plane side, never in the browser. What can I use today? The Apache-2.0 control plane, CLI, gateway, local quickstart, SDK source, and provider integrations are public today. The hosted web console remains a private preview. InferCrane Cloud private preview Want production inference without operating the platform yourself? The open-source inference operations platform is available now. Join the waitlist for InferCrane Cloud and hands-on launch support. Get launch updates and request private-preview access. No spam. Confirm by email. Read our privacy notice.