AI News HubLIVE
站內改寫4 分鐘閱讀

待翻譯:Tinker: GLM 5.3 Fine-Tuning

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Clock Cycles & Pipelining Session Metrics OpenAI-Compatible API Anthropic-Compatible API API Reference Storage Contributing 200: Core Concepts 300: Cookbook Abstractions 400: Advanced 500: Deployment Changelog Models &…

來源Hacker News AI作者: tosh

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Clock Cycles & Pipelining Session Metrics OpenAI-Compatible API Anthropic-Compatible API API Reference Storage Contributing 200: Core Concepts 300: Cookbook Abstractions 400: Advanced 500: Deployment Changelog Models & Pricing All prices are per million tokens. Checkpoint storage is charged at $0.10 per GB per month. Training We provide an 80% discount on cached prefill tokens. All Types All Types Base Reasoning Hybrid Vision All Architectures All Architectures Dense MoE All Sizes All Sizes Compact Small Medium Large Model Tinker ID Context Size Arch Type PrefillCached: 80% discount Sample Train InklingLimited-time 50% discountthinkingmachines/Inkling64KLargeMoEHybrid + Audio + Vision$3.74 $1.87$0.374 (cached)$9.36 $4.68$11.22 $5.61 Inkling (256K)Limited-time 50% discountthinkingmachines/Inkling:peft:262144256KLargeMoEHybrid + Audio + Vision$7.48 $3.74$0.748 (cached)$18.72 $9.36$22.46 $11.23 Inkling-SmallLimited-time 50% discountthinkingmachines/Inkling-Small64KLargeMoEHybrid + Audio + Vision$1.16 $0.58$0.116 (cached)$2.88 $1.44$3.46 $1.73 Inkling-Small (256K)Limited-time 50% discountthinkingmachines/Inkling-Small:peft:262144256KLargeMoEHybrid + Audio + Vision$2.32 $1.16$0.232 (cached)$5.78 $2.89$6.94 $3.47 Nemotron-3.5-Lightning-30B-A3BLimited-time 50% discountnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF1664KMediumMoEHybrid$0.39 $0.195$0.039 (cached)$0.99 $0.495$0.88 $0.44 Nemotron-3.5-Lightning-30B-A3B (256K)Limited-time 50% discountnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16:peft:262144256KMediumMoEHybrid$0.52 $0.26$0.052 (cached)$1.32 $0.66$1.60 $0.80 Nemotron-3-Ultra-550B-A55BLimited-time 50% discountnvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF1664KLargeMoEHybrid$4.98 $2.49$0.498 (cached)$12.45 $6.225$10.956 $5.478 Nemotron-3-Ultra-550B-A55B (256K)Limited-time 50% discountnvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16:peft:262144256KLargeMoEHybrid$6.64 $3.32$0.664 (cached)$16.60 $8.30$19.92 $9.96 Nemotron-3-Super-120B-A12BLimited-time 50% discountnvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF1664KMediumMoEHybrid$1.14 $0.57$0.114 (cached)$2.88 $1.44$2.552 $1.276 Nemotron-3-Super-120B-A12B (256K)Limited-time 50% discountnvidia/NVIDIA-Nemotron-3-Super-120B-A12B-BF16:peft:262144256KMediumMoEHybrid$1.52 $0.76$0.152 (cached)$3.84 $1.92$4.64 $2.32 Nemotron-3-Nano-30B-A3BLimited-time 50% discountnvidia/NVIDIA-Nemotron-3-Nano-30B-A3B-BF1664KMediumMoEHybrid$0.39 $0.195$0.039 (cached)$0.99 $0.495$0.88 $0.44 GLM-5.3 (256K)zai-org/GLM-5.3:peft:262144256KLargeMoEReasoning$4.86$0.972 (cached)$12.15$14.58 Kimi-K2.6moonshotai/Kimi-K2.632KLargeMoEHybrid + Vision$2.205$0.441 (cached)$5.49$4.84 Kimi-K2.6 (128K)moonshotai/Kimi-K2.6:peft:131072128KLargeMoEHybrid + Vision$5.15$1.03 (cached)$12.81$15.40 Qwen3.8-27BQwen/Qwen3.8-27B64KMediumDenseHybrid + Vision$1.86$0.372 (cached)$5.595$4.103 Qwen3.8-27B (256K)Qwen/Qwen3.8-27B:peft:262144256KMediumDenseHybrid + Vision$2.48$0.496 (cached)$7.46$7.46 Qwen3.6-35B-A3BQwen/Qwen3.6-35B-A3B64KMediumMoEHybrid + Vision$0.54$0.108 (cached)$1.335$1.177 Qwen3.6-27BRetiring September 2Qwen/Qwen3.6-27B64KMediumDenseHybrid + Vision$1.86$0.372 (cached)$5.595$4.103 Qwen3.5-397B-A17BQwen/Qwen3.5-397B-A17B64KLargeMoEHybrid + Vision$3.00$0.60 (cached)$7.50$6.60 Qwen3.5-397B-A17B (256K)Qwen/Qwen3.5-397B-A17B:peft:262144256KLargeMoEHybrid + Vision$4.00$0.80 (cached)$10.00$12.00 Qwen3.5-35B-A3B-BaseQwen/Qwen3.5-35B-A3B-Base64KMediumMoEBase$0.54$0.108 (cached)$1.335$1.177 Qwen3.5-9BQwen/Qwen3.5-9B64KSmallDenseHybrid + Vision$0.66$0.132 (cached)$1.995$1.463 Qwen3.5-9B-BaseQwen/Qwen3.5-9B-Base64KSmallDenseBase$0.66$0.132 (cached)$1.995$1.463 Qwen3.5-4BQwen/Qwen3.5-4B64KCompactDenseHybrid + Vision$0.33$0.066 (cached)$1.005$0.737 Qwen3-8BQwen/Qwen3-8B32KSmallDenseHybrid$0.195$0.039 (cached)$0.60$0.44 GPT-OSS-120Bopenai/gpt-oss-120b32KMediumMoEReasoning$0.33$0.066 (cached)$0.84$0.737 GPT-OSS-120B (128K)openai/gpt-oss-120b:peft:131072128KMediumMoEReasoning$0.78$0.156 (cached)$1.94$2.33 GPT-OSS-20Bopenai/gpt-oss-20b32KSmallMoEReasoning$0.18$0.036 (cached)$0.45$0.396 DeepSeek-V3.1deepseek-ai/DeepSeek-V3.132KLargeMoEHybrid$1.695$0.339 (cached)$4.215$3.718 Serverless Inference (Beta) Serverless inference is currently in beta and available for Inkling and Inkling-Small only. We do not recommend it for intensive production use until it is out of beta. If you're interested in production use, email us at [email protected] to join the waitlist. Please include which models you need, your expected volume and latency requirements, and your use case. Model Tinker ID Context Prefill (Input) Sample (Output) Inkling-Smallthinkingmachines/Inkling-Small:peft:262144:sampling-nvfp4256K$0.30$0.06 (cached)$1.20 Inklingthinkingmachines/Inkling:peft:262144:sampling-nvfp4256K$1.00$0.17 (cached)$4.05 Pricing Terms Prefill: Processing input/prompt tokens (forward pass only) Cached prefill: The smaller price under each prefill price; applies to input tokens that hit the prompt cache (80% off) Sample: Generating output tokens (forward pass + sampling) Train: Forward and backward pass for gradient computation Context: Maximum sequence length. Models with :peft: suffix support extended context at higher prices. Tinker ID: The exact string to pass to create_lora_training_client(base_model=...) or create_sampling_client(base_model=...) MoE models are priced by active parameters, making them significantly more cost-effective than dense models of similar quality. Model Types Base: Raw pretrained models with no chat or instruction tuning. Best for post-training research or running the full post-training pipeline yourself. Reasoning: Always produce chain-of-thought before their answer. Highest intelligence, higher latency and token cost. Hybrid: Run in both thinking and non-thinking modes. They reason by default, but chain-of-thought can be disabled via a renderer or argument for faster, cheaper direct answers. Vision: Vision-language models that accept images alongside text. Shown as a + Vision suffix on the underlying type (for example, Hybrid + Vision). Audio: Models that accept audio alongside text. Shown as a + Audio suffix on the underlying type. Architecture is either Dense (all parameters active per token) or MoE (mixture-of-experts, only a subset of parameters active per token). MoE models are highlighted in amber. Choosing a Model Cost-effective: Use MoE models (highlighted in amber) Research/post-training: Use Base models Task-specific fine-tuning: Start with a Hybrid model Low latency: Use a Hybrid model with chain-of-thought disabled High intelligence: Use Reasoning or Hybrid models (chain-of-thought) Vision tasks: Use models with Vision in the type Machine-Readable Pricing If you want to use this data programmatically (cost estimation, model pickers, dashboards), don't scrape the tables above. The data behind them is published as JSON alongside this page, and those files are the stable interface for scripts (the tables' HTML is presentational and may change): models.json: the Training table. One object per model with name, tinker_id, context, size, arch, type, url, and per-million-token prices prefill, cached_prefill, sample, train (strings like "$0.374"). Temporarily discounted models also carry original_* price fields and a note. serverless.json: the Serverless Inference table, with name, tinker_id, context, url, input, cached_input, output. For example: import json, urllib.request url = "https://tinker-docs.thinkingmachines.ai/tinker/models.json" models = json.load(urllib.request.urlopen(url)) for m in models: print(m["tinker_id"], m["train"]) Retired Models These models have been retired and can no longer be used for training or inference, grouped by retirement date. See Model deprecations for the recommended replacement for each. " subsection above the others (most recent first) and list the models grouped by family. --> July 12, 2026 Kimi: Kimi-K2.5 June 12, 2026 Qwen: Qwen3-235B-A22B-Instruct-2507, Qwen3-VL-235B-A22B-Instruct, Qwen3.5-35B-A3B, Qwen3.5-27B, Qwen3-32B, Qwen3-30B-A3B, Qwen3-30B-A3B-Instruct-2507, Qwen3-VL-30B-A3B-Instruct, Qwen3-30B-A3B-Base, Qwen3-8B-Base, Qwen3-4B-Instruct-2507 Llama: Llama-3.3-70B-Instruct, Llama-3.1-70B, Llama-3.1-8B, Llama-3.1-8B-Instruct, Llama-3.2-3B, Llama-3.2-1B DeepSeek: DeepSeek-V3.1-Base Kimi: Kimi-K2-Thinking