AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。
All blog posts Inference Published 9/9/2026 The Open Source AI Stack Authors Hassan El Mghari Table of contents 40+ Models Chosen for Production...40+ Models Chosen for Production...40+ Models Chosen for Production... As the quality of open source models have bridged the gap with closed source models, a lot of developers and organizations are looking to move to open source models for more ownership, control and economics. This post is a deep dive into the open model AI stack that developers need to consider as they move from closed to open source. Using open models for agentic software development does not require learning how to train models, buying a rack full of GPUs, or becoming an expert in machine learning. From the perspective of an application developer, the stack is surprisingly familiar to using closed-source models. You have a model that answers prompts and a harness that manages the interaction between you and that model. If you already know how to use Claude Code, you’re much closer to using open models than you probably think. The MIGHT Stack The stack can be broken down into these separate parts: Model: The component that interprets your request and decides what to do. Inference: The infrastructure & inference provider where the model actually runs. Gateways and routers: The layer that decides which model or provider handles each request, balancing cost, speed, and capability. Harness: The application that manages the conversation, gives the model access to tools, and connects it to your codebase. Tools (Skills and MCP): Knowledge that tells the model how to do specific tasks, including giving models and the harness access to relevant context These layers are all independent from one another, which opens the door for technical decisions at each layer that better suit your development workflow. In turn, this allows you to experiment with new models as soon as they are released. New models appear constantly. Some are faster. Some are cheaper. Some are unusually good at a particular kind of work. If switching models takes a few minutes instead of rebuilding your workflow, you can actually experiment with them. In this post we will explore each layer of the stack, explain what it does, and look into how it can be customized. Let’s start with the model layer. Models A model takes your prompt and generates a response by predicting the most likely next tokens based on its training data. In our stack, it acts as the intelligence layer and it’s responsible for reasoning, decision-making, and deciding what changes to make in your codebase. Models come in a range of sizes, and in general, larger models tend to have greater capability, better reasoning, and more reliable performance on complex tasks. Most leading open models are now Mixture-of-Expert (MoE) models that contain many specialized “experts”, but only activate a small subset of them for each token generation. This allows them to have larger parameter counts but only activate a smaller of parameters and thus need less compute to run. Large models Large models are typically defined by the number of parameters they contain and the amount of compute used to train them. These models have an extraordinary capacity to recognize patterns, relationships, and abstractions from their training data. In practice, “large” usually also implies that the model was trained on more data, for longer, and with significantly more compute resources. This extra capacity translates into a few important benefits. Large models tend to be better at multi-step reasoning, where they need to keep track of several constraints at once and make decisions that depend on earlier parts of the problem. This makes them excellent choices when given an ambiguous or under-specified task since they can draw on a broader range of learned patterns to fill in missing details. A good example of a large open model is Kimi K3, which has 1.8T total parameters & 104B active parameters. Reach for large models like Kimi K3 when you need to perform complex tasks, such as: Refactoring an existing authentication system Upgrading your codebase to a new framework Reviewing pull requests Understanding why an SQL database all of a sudden became slow An advantage of large models is their robustness across different types of work. They can switch between writing code, explaining systems, debugging issues, and planning changes without needing tightly scoped instructions. This makes them especially useful in agent-style workflows where the model has to decide what to do next rather than simply follow a single instruction. Large models also make better use of longer conversations. When they are given many files, logs, or pieces of information at once, they are able to maintain coherence and connect relevant details across the entire input. This might lead you to believe that larger models are always better, but in practice there’s a tradeoff between large and small models. Next, we’ll look at some reasons to choose smalls model over larger ones. Small models The difference between small and large is less about quality and more about how much ambiguity they can comfortably handle. When a task is clearly defined and tightly scoped, small models can perform on par with much larger ones. If you remove ambiguity by being explicit about what you want, they become extremely effective. Small models excel at well-specified work since they don’t need to guess your architecture, infer hidden requirements, or explore multiple possible interpretations. They just execute the instruction as given. An example of a small open model is GLM 5.3 Flash which has 320B total parameters & 18B active parameters. To compare it to Kimi K3, it’s ~6 times smaller and ~20 times cheaper. These models are surprisingly good when the task is narrow and well-specified, for example: Update this function to accept another option. Write tests for this file. Explain a specific error. Review this 50-line function for bugs. Rename this API and update its callers. There is not much ambiguity in these tasks. The model does not need to build a detailed understanding of your entire codebase or decide among several different architecture approaches. The most important advantage of small models is that they are significantly faster and cheaper to run. They require less compute to generate each token. In practice this means lower latency responses and a much lower cost per request. For many day-to-day coding tasks, this speed difference is immediately noticeable since the model responds quickly, iterations happen faster, and you can afford to run it repeatedly without worrying about cost. Rather than thinking of bigger models as being better than smaller models, it’s better to think of these models as two different types of tools in a toolbox. Models as tools A larger model might solve a problem more reliably, but it might also be several times slower or more expensive. If a smaller model can make the same one-file change correctly, there is little reason to reach for the larger one. Over time, start treating model selection more like choosing a tool rather than picking a winner. Start with a handful of both large and small models, and learn their capabilities and limitations through repeated use. You’ll quickly form an opinion about what tasks in your codebase can be delegated to differently size models. Next, we'll look at the best places to find models. Picking the model New models are released weekly and there are too many models to evaluate all of them yourself. Leaderboards such as The Open Frontier and Artificial Analysis are useful for discovering what is available and getting a rough sense of how models compare across intelligence, coding ability, speed, and price. Do not spend too much time trying to identify the single best model on a leaderboard. Benchmarks compress a lot of behavior into one score, while your actual workload is much more specific. A model that ranks slightly lower overall might be excellent at the kind of coding work you do every day. For reference, the most popular open models in the space right now are GLM 5.3 Flash, DeepSeek V4 Flash, Kimi K3, and MiniMax M3. These change very rapidly though. After you’ve selected a few models to try out, the next step is to find a provider that hosts them. Inference providers Because open models are available outside a single ecosystem, you get to choose where they run. The easiest way to get started is with cloud inference providers and cloud gateways. The way these providers work is that you send them an API request and they run the model on their GPUs. You pay for the input you send to the model and the output from the model, measured in tokens. This makes cloud providers ideal for experimentation. If you want to try a new model, you can create an API key, specify a model name, and start sending requests. Providers such as Together AI and others offer large catalogs, so a single account can give you access to many different large and small models. Two providers running the same model should generally produce similar results, especially when you use the same model version and sampling settings. Performance, price, latency, and API features may differ, but the underlying model is still the same. This separation means that you get to choose a model because you like its behavior, then a provider or gateway based on price or performance. Gateways and routers Models are not uniformly hosted by all inference providers, which unfortunately means you cannot use a single inference provider for every model. Gateways and routers solve this by giving you access to many different models across multiple inference providers. This is especially useful if you want to route between closed and open models. A gateway sits in front of multiple inference providers and aggregates them behind a single API, letting you route requests across different backends, compare pricing and latency, and switch models without changing your code. It acts as a translation and routing layer between you and the underlying inference providers. Two popular cloud gateways worth exploring are https://openrouter.ai/ and https://vercel.com/ai-gateway. You can also run your own router locally, or on your own server, with tools like https://www.litellm.ai/. These routers give you a single API endpoint that lets you route between different accounts that you’ve already set up with multiple inference providers. Once you’ve selected an inference provider or cloud gateway, the next step is to set up your harness. Harnesses The harness is the part of the stack you interact with. It’s often a program running on your computer that sits between you, your codebase, and the model. It maintains the conversation and gives the model tools it can use to interact with the real world. Suppose you ask: Where in this codebase is authentication handled? The harness sends that question to the model. The model might decide the first thing it needs to do to answer that question is search the codebase for terms such as auth, session, or login. It responds by asking the harness to run those searches. The harness runs the search commands on your computer and sends the results back to the model. Based on those results, the model asks the harness to read the code from several files. The harness adds the contents of those files to the conversation with the model. After a few rounds of this, the model has enough information to answer your authentication question. It can answer with all of the files and lines in your codebase that contain auth code. The important part is that the model is not searching your filesystem or executing shell commands. The model decides what should happen and the harness is the piece that makes it happen. This means the quality of a coding agent depends on more than just the model. A g [truncated for AI cost control]