跳到主要內容
AI News HubLIVE
站內改寫6 分鐘閱讀

待翻譯:Making the leap to specialized intelligence

文章摘要

AI 服務暫時不可用,以下為來源摘要,待恢復後補全翻譯:Making the leap to specialized intelligence Join us for our inaugural conference, Forge 2026 Blog Making The Leap To Specialized Intelligence Making the leap to specialized intelligence PUBLISHED 9/10/2026 Table of Cont…

待翻譯:Making the leap to specialized intelligence
報告錯誤

更正渠道尚未開通,可先複製下方文章資訊留存。

查看更正說明
直接讀正文

AI 服務暫時不可用,以下為來源正文,待恢復後補全翻譯。

Making the leap to specialized intelligence Join us for our inaugural conference, Forge 2026 Blog Making The Leap To Specialized Intelligence Making the leap to specialized intelligence PUBLISHED 9/10/2026 Table of Contents TL;DR The road to specialized intelligence Stage 1: Renting the closed frontier Stage 2: AI engineering - prompts, context & harnesses, oh my! Stage 3: Adding open models into the mix Stage 4: Training - owning your specialized intelligence "Teach the model our taxonomy" "Write the way we write" "The big model works, we just can't afford it at scale" "Make the agent actually finish the job" "Our domain changes often" Taking the first step to specialized intelligence Explore training options Fireworks offers three paths to train for every level of expertise Learn more This is the first piece in a series on specialized intelligence for engineers who are dabbling in training, or are curious about what it would take. TL;DR •Training your own models with specialized intelligence means owning what makes your company unique •The road to specialized intelligence can start with renting the closed frontier, but inevitably leads to training open models •Common patterns emerge as to why companies turn to training their own models The road to specialized intelligence Teams decide to train their own models for a handful of reasons: •Fear that a closed frontier lab will move into their domain •A quality bar the general models don't clear for their vertical •An API bill that's too high for the intelligence they actually need Underneath all of these is the same instinct: to own the specialized intelligence that defines their business. This piece is about the road that gets you there. You start by renting the most powerful (usually the most expensive) closed LLM available to you, engineering specialized harnesses, systems, and components around that model to improve its performance on your task. Once those components max out, you migrate to cheaper, yet still effective, open models to cut down on cost while maximizing utility. Finally, if the problem is truly yours, you decide to own the intelligence outright by training a custom model. Stage 1: Renting the closed frontier Almost everyone starts with a closed frontier model behind an API. It makes sense: those providers spend enormous sums marketing their models everywhere they can, trying to get you on board. These models are powerful and because someone else carries the cost of training and serving them, you are paying by usage through token costs. The downsides are deep but here are three main ones you might run into: The model is far too general and demands prompting on top to be able to do anything in a specialized way. When it works, it works, but the model is molded by someone else, for everyone else. A closed API is a rental in every sense. You don't control the model weights, the price, the latency, the deprecation schedule, and in most cases, you do NOT control your data. Your token premiums and the enterprise subscriptions you hand out to your teams are what pay for all of it: the infrastructure you never touch, the marketing campaigns, the consultants advising on those campaigns, and the R&D that keeps the lab at the frontier. You are subsidizing a machine you don't own and can't steer. After all that work they put into their models, closed frontier LLMs do indeed come out powerful, but they also come off the shelf far too generalized and they need to be fed relevant context to be useful day to day. Enter AI engineering. Stage 2: AI engineering - prompts, context & harnesses, oh my! The last few years of AI engineering have followed a clear escalation, each layer added to squeeze more out of a model we were renting off the shelf. First came prompt engineering: few-shot examples and chain-of-thought showed that changing how we asked induced stronger, more consistent behavior. That turned into context engineering as we piped more and more into the context window, usually for retrieval-augmented generation (RAG), to keep the model current with shifting information. And once prompts and context were handled, agents arrived and filled that same window with loops of tool calls, so we turned to harness engineering: conversation compaction, tool optimization, MCP, and more to treat the context window as a budget rather than a bucket. The whole arc has been humans picking up the engineering slack of a model that never changes. This is the stage most serious teams live in for a long time, and it's the first time you start to truly feel like you’re owning your AI. Prompting, tool use, and RAG can carry you remarkably far; but there's a ceiling. Every model in your harness is still the same model it was on the shelf. Your prompts and harnesses guide the models but they don't teach them anything. And everything the harness knows about your business, it re-explains on every single request in tokens you pay for, in a prompt you re-paste into a model that might be deprecated next week. Every AI system is really being optimized along three axes: quality (is it good enough?), cost (can you afford it at scale?), and latency (is it fast enough?). AI engineering is the effort to push all three without touching the model itself and it can take you a long way before you hit the ceiling on any of them. Addressing the ceilings of performance is difficult, which is why most people take aim at optimizing cost and latency first and one of the most effective ways of doing this is by mixing in open models. Stage 3: Adding open models into the mix These days, frontier open models tend to get the job done at a fraction of the cost of a frontier closed model, and depending on the task, the quality gap is small or nonexistent. This stage is defined by unit economics and control, not ideology. By this stage, you now want to decide where the model runs, how fast it responds, what happens to your data, and what you pay per token and per task. To pick the right model (open or closed), you'll either lean on reported benchmarks or, better yet, evaluate candidates on your own task. Benchmarks aren't always representative of the work you're actually doing. Maybe your fintech use case is more niche than the benchmark's examples, or your legal questions span shifting philosophies across a dozen domains. Benchmarks are a decent way to shortlist; your own evals (systematic homegrown tests for AI systems on your specific tasks) will always be the best indicator of performance. BenchmarkKimi K3 (max reasoning)GPT 5.6 Sol (max reasoning) OSWorld 2.058.362.6 UIPad87.787.7 Let’s say you are building agents that need to navigate UIs in order to extract information, interact with elements on the screen, or simply answer questions. This is called computer use and there are benchmarks for this. For example, if you look at the leaderboard for OSWorld 2.0, a benchmark for multimodal computer use (meaning AIs are visually looking at a browser and actually navigating it to solve a task), GPT 5.6 Sol beat Kimi K3 by only a few percentage points. You could simply call it a tie and go with the cheaper Kimi K3, but there’s always a level deeper we could go. Let’s take another dataset for computer use as an example. It’s called MacPaw UIPad which has over 1,000 examples of screenshots paired with questions and answers to those questions. This dataset is not a common benchmark, but exists as a publicly available test for computer use. When I ran K3 and Sol on this dataset, they both got the same grade on UIPad but Sol cost $115 to run against the dataset whereas K3 cost $56 (over 2x cheaper). So far, it’s seemingly the same result as OSWorld. Both models did just as well as the other, but one was cheaper: decision made. But there’s something interesting about the dataset: it’s split up into four categories, and if we dig into each of those categories, something interesting appears: yes/no questions - K3 wins questions where numbers are the answer - K3 wins questions where there’s a natural language answer - a tie coordinate questions where the AI must draw a box around the item being asked for - Sol wins K3 and Sol tie at 87.7 overall on the UIPad test set, but the category split is lopsided: K3 wins or ties three of four while Sol takes coordinates by 13 points. That gap is an argument for routing. Knowing ahead of time that K3 ties or wins on 3/4 categories while being cheaper means you could route coordinate tasks to Sol and everything else to K3 and end up with a stronger and cheaper system overall. Engineering the strongest possible system with the freedom to experiment with the most powerful AI models on the planet is intoxicating indeed, but again, there’s a ceiling. AI engineering is a bandaid for owned intelligence, not the real solution. Stage 4: Training - owning your specialized intelligence At some point your business logic, your taxonomy, your in-house style, your customers' quirks are the product. This is where those three motivations, plus the ownership instinct beneath them, stop being reasons to consider training and start being reasons to do it. Training (aka fine-tuning, the act of changing a model’s intelligence in place) is how you stop re-specifying those on every request and move them into model weights you control. The model stops being a generic tool you rent by the token and becomes an asset you own. In practice, training models shows up as five common patterns. By operating on the premise that AI models off-the-shelf already have a base level of intelligence, these patterns work to mold a model’s intelligence to be specialized to your data and workflows. "Teach the model our taxonomy" This is the broadest, most encompassing pattern. Generally speaking, we need our model to go from a generic interpretation of a task (often tested via public benchmarks) to a more specific version of the task (tested by an eval on an internal test set). Classification and extraction are two classic examples. Imagine bucketing a support ticket into 1 of 17 topics (that’s classification), escalating to the right team (also classification), reading a receipt for expense tracking (extraction). These sound simple, but the model doesn't know your decision boundaries, or the nuance between two topics that only you understand. You can prompt and few-shot your heart out, but you're paying for those tokens forever and still gambling on the gray areas. Give the model a few hundred to a few thousand examples of "input → the label we actually use," and it stops guessing at your taxonomy and starts knowing it. Our UIPad example also fits in this pattern. It’s neither extraction nor classification but we are trying to get our model to understand our specific version of a task. So let’s put this to the test (pun intended). Here’s the plan: Split the UIPad dataset into a training / testing set This way we can be more confident that the model can generalize its learnings from the training set to unseen data on the test set Run Kimi K3 and GPT 5.6 Sol on the test set to get a baseline score Note beforehand, I was showing you scores on the whole dataset, not split up so we need to do this again to get a fair comparison Train K3 on the training set and not the test set I used the Fireworks Serverless Training API I pointed my coding agent to our cookbook’s training skill file and let it do the rest Run the new trained model against the same test set to get a final score Base K3 trailed Sol on the held-out test; after roughly three hours of training, tuned K3 clears it. Et voilà! Our tuned K3 model is now better than GPT 5.6 Sol after only 3 hours of training. This is just one example of how to combine unique data with models off the shelf to create models with true specialized intelligence. "Write the way we write" If you’re using AI to write things like marketing briefs [truncated for AI cost control]

展開要點與分析

文章情報

工程師進階

要點

  • AI 服務暫時不可用,系統已先保留來源內容與降級元數據。
  • Making the leap to specialized intelligence Join us for our inaugural conference, Forge 2026 Blog Making The Leap To Specialized Intelligence Making the leap to specialized intell…

技術影響

可能影響 Agent 架構、工具調用、工作流自動化和產品集成。

要點與分析由自動化流程生成,可能有誤,請結合原始來源核實。