AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Take that, OpenAI! Anthropic! Chinese AI models have surpassed their U.S. counterparts in token consumption on OpenRouter. You might think U.S. AI companies dictate the AI economy. You’d be wrong. According to dat…
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Take that, OpenAI! Anthropic! Chinese AI models have surpassed their U.S. counterparts in token consumption on OpenRouter. You might think U.S. AI companies dictate the AI economy…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Today on Decoder, I’m talking to Verge senior AI reporter Hayden Field about some pure Decoder bait: the seemingly-endless org chart changes at OpenAI, and how all of them seem to consolidate power under cofounder Greg Brockman, the company’s president. While Sam Altman is the CEO and still OpenAI’s most public face, Brockman has amassed enormous power and influence within the top ranks of the company as other senior leaders have left in rapid succession these past few months. Verge subscribers, don’t forget you get exclusive access to ad-free Decoder wherever you get your podcasts. Head here. Not a subscriber? You can sign up here. Brockman now oversees the company’s entire consumer and enterprise product teams, including ChatGPT, Codex, and its major infrastructure build out. As Hayden recently reported, he is effectively now the day-to-day operational leader of OpenAI. This will have major effects on OpenAI as a company and its product strategy, at a time when it continues to cede ground to Anthropic in the enterprise and is preparing for a historic IPO. This is all against the backdrop of needing to turn a profit in the next few years and the company’s huge ambitions to replace both Google Search and the iPhone in the consumer market. So I wanted Hayden to break down for me what’s going on at OpenAI, and the increasingly important role Greg Brockman will play in its future. Okay: Verge senior AI reporter Hayden Field on Greg Brockman’s consolidation of power at OpenAI. Here we go. This interview has been lightly edited for length and clarity. Hayden Field, you’re The Verge‘s senior AI reporter. Welcome back to Decoder. Thanks. It’s great to be here. Always, always chaos when you’re here, Hayden. Absolutely. There’s never a calm week. If there’s a calm week coming up, I know something even crazier is coming the next week. We should just rename the show The Real Housewives of AI. [Laughs] Honestly, that would be fitting. Straight up, that’s what we should do. A lot of personalities, a lot of feelings, a lot of relationships that people have really come to value over a long period of time. And a lot of lore. A lot of lore that goes between all these people for years and years and years. It’s crazy. And now they’re even putting out profiles on some of their spouses. The circles and the people themselves are really interesting. Today the drama is about OpenAI and specifically Greg Brockman, who seems to be consolidating even more power at OpenAI. You just wrote a long story about this. It seems very clear that Greg is emerging as the central decision maker at OpenAI. Describe what’s going on. Greg has obviously been pretty influential at OpenAI for a really long time, but what’s different now is that he has so much control over the day-to-day operations in a way he didn’t before. He was always a cofounder. He was always heavily involved. It was him, Ilya Sutskever, and Sam Altman for a long time, all in these email threads that came out during the Musk v. Altman trial. We saw them all talking about the future of OpenAI, what it should look like, and all their strange dynamics with Elon Musk. But Brockman was in a big-picture role before, and now he’s increasingly taking on so much power in the day-to-day. Altman has been focusing more on the big picture, the IPO stuff, and the direction of the company at large. Meanwhile, Brockman is amassing all this power day-to-day as other executives leave. He’s been in control now of Codex, of its enterprise business, the consumer side. There are four arms of the company right now and he’s in control of all of them. It’s been really interesting to see how that’s happened over the last few months. Greg is a character in the AI story. Certainly inside of OpenAI, he’s been a central figure in a lot of the company’s drama over time. So if you’ve been paying attention to the AI industry, you’re familiar with him. But if someone haven’t, explain who Greg is and why he’s such a notable figure both in AI and at OpenAI. He has been on my radar for seven years now. He’s been on other people’s for even longer. He dropped out of MIT to join Stripe way back in the day. He was Stripe’s CTO during its really early explosive growth phase. He left in 2015 to cofound OpenAI. That’s how he came on the AI radar. Ever since then, he’s been a really influential figure. He, Ilya Sutskever, and Sam Altman were all heavily involved in the drama with Elon Musk really early on, raising money, trying to lure people over to OpenAI, hire top talent, and figure out who was going to control AGI in the event they ever made it. There’s a lot going on there. In the years following, Mira Murati joined and Brockman, Sutskever, Murati, and Altman became the four main players here. These were the most talked-about people at OpenAI. You saw them in the headlines the most. They had the most power at the company. And then when the board coup happened in 2023, that’s when Brockman flew onto the scene in a new way. He had been behind the scenes a little bit, under the radar. He was not really doing that many public interviews. I interviewed him in 2020, but it’s not like he was a talking head. You weren’t seeing that much of him unless you were really watching the company and the industry. But when Sam Altman got fired by the board, Brockman was so incensed that he quit immediately. He and Altman were going to start their own thing, or were potentially going to lead a department at Microsoft. That’s when he was making waves in a new way because he was hellbent on the drama. Brockman said, “Okay, if Sam goes, I go.” That’s when he became Sam’s right hand in a new way. Before that, he was just one of the many execs. They were close, but the board coup made them closer than ever. That’s also what made Sam know that he could trust him in a new way and thereby later greenlight him getting more power. This was a big bet in the politics of the organization. Mira was going to become the new CEO for five minutes. Then there was going to be another new CEO and it was unclear who would stay and who would go. Justifiably, Altman did not know who he could trust in that moment, except Greg Brockman was ride or die. He said, “I’m leaving. I’m going.” The five minutes that they were going to start a new division of Microsoft were some of the most hilarious five minutes in tech history. But Brockman had basically made it clear he was ride or die for Sam. Over time, that bet has paid off. That move has paid off. Even during that board coup moment, it was not clear that Sam would ever come back. I remember I was backpacking in Patagonia at the time and I didn’t have a laptop, but I wrote six articles on my phone. It was a crazy time. It seemed like Brockman and Altman were going to be fine, but just be on their own. Of course, all the hundreds of employees signing a letter that they would leave if Sam was not reinstated is what got the ball rolling on him coming back and then pulling an UNO reverse on all the board members who had voted to oust him, except one. That was a big moment where he didn’t know who he could trust. Some of the people who he had trusted the most were the ones who voted for him to be fired. That went a long way with him that Brockman quit as well. Ever since then, we’ve seen Brockman positioned in the public eye a little bit more. He’s been in the press more. This summer is when he really amassed a lot more power after a couple other executives left, and he’s now in control of quite a lot of pillars of the company. What’s really striking is, yes, we can talk about 2023 and the board coup and Sam going and coming back and who was going to stay or go then, but actually this year has been more dramatic in terms of executive departures from OpenAI. Just since April, the list is staggering. I’m just going to run down it: Bill Peebles who ran Sora; Kevin Weil, who was head of product at Instagram and then VP of OpenAI’s science arm; Fidji Simo, who was supposed to be head of product and was at one point the AGI chief, left on health leave and then just never came back; Kate Rouch, the Chief Marketing Officer; Srinivas Narayanan, the CTO of B2B; Brad Lightcap, who was the former CEO and the head of special projects; and then just recently Denise Dresser, the Chief Revenue Officer, who is an important character if your company’s about to IPO. That is a lot of people, and it seems like Brockman just took all of their roles as all of those people left. Is that what’s happening here? More or less. Some of them definitely affected his position less than others. For some of them, like CMO Kate Rouch, leaving didn’t really give him that much more power. But what was interesting is Simo, the ex-chief of AI at the company, is the one whose absence left a huge gap for Brockman to fill. He took over a ton of different aspects. As you just alluded to, a lot of the people who had just left the company had also recently changed roles. Through one way or another, a lot of these people leaving created a power vacuum that Brockman could then capitalize on. It’s not like he had this master plan of amassing power, but either way, he ended up with a lot of power. What is interesting is when Brad Lightcap transitioned roles, for example, one of these other people took over his COO responsibilities. Guess who it was? Dresser. Then she leaves. Now some of those responsibilities fall to Brockman. It didn’t all happen cut and dry, but in one way or another, people switching roles and then leaving the company, he ended up with a ton more power. I can’t even calculate the ratio really of how much power he ended up with. Now he’s in control — to give you some context — of the entire product strategy of the company, which is obviously incredibly important when you’re about to IPO and you’re getting a lot of pressure to turn a profit. He’s in charge of the company’s entire scaling arm, as well as core product and platform; all of consumer, which includes health, commerce, personal finance, ChatGPT, and other consumer-facing stuff; and core infrastructure, ads, data science, and growth. What isn’t he in charge of, really? You know what I mean? I want to talk about Fidji Simo for one more second here. She was the former CEO of Instacart, but before that she was a really high-ranking executive at Meta. She ran the Facebook app. She was in charge of a lot of monetization. When she came to OpenAI, it felt like her role was to turn the consumer version of ChatGPT into a product. There were lots and lots of people leaving Meta to go to OpenAI for a while, such that basically we had reported that every all hands at Meta was about, “What are you going to do about OpenAI? They’re taking all of our people.” You just saw this exodus of people from Meta going to OpenAI in various ways, led by Fidji Simo, in a moment where it felt like what they wanted to do was take on Google as a big consumer product supported by ads. It’s funny, Altman actually copped to that ambition on the David Senra’s show last week — here he is, giving Peter Thiel credit for saying OpenAI should compete directly with Google: Sam Altman: He’s like, ‘The power of this is the power of the Google text box, it’s a text box you can type anything into and it does the right thing. Clearly, the empty text box worked for Google so why don’t you just double down on that?’ That’s pretty straightforward and that’s what they did. Altman says they went super hard at it — being the empty text box, taking the Google business model, which is one of the best business models in internet history. That was a big bet, and they hired Fidji Simo to go run that bet. But she left, and that does not seem to be the emphasis anymore. The emphasis is on Codex, on enterprise, on competing with Anthropic. What’s left of that? Was that just a misfire and she had the wrong ideas? Is it still the plan? This seems like the question about OpenAI. Some of it has fallen by the wayside, which we’ve seen for a bunch of stuff that OpenAI has tried in the past year or two. They vowed to stop doing all their side quests, which is ironically something that Simo herself wrote in an internal memo. Several episodes ago, you were on and we just talked about the series of code reds at OpenAI. Right now they’re focused on the key revenue drivers, which are enterprise and coding. That’s really what they’re all-in on. They’re also, of course, really focused on building a “super app,“ which is supposed to incorporate all this stuff together, and making that super app better. Besides that, they’re trying to cut the fat and operate as a leaner business, which has been helped by some of these executive salaries being cut. They’re trying to look better on their balance sheet when they IPO. There’s a lot of pressure to do that, especially when SpaceX just IPOed at a crazy valuation and Anthropic is apparently gunning to exceed that. We’ll see if that actually happens. When they changed Simo’s title from CEO of applications to CEO of AGI deployment, that was the most Decoder thing of all. I looked at their org chart changing and I thought, “This isn’t going to last.” Over and over and over again, they keep reassigning people. And then as you say, they leave and we see one person take over those roles and consolidate power. We’ll come to the IPO in a second, but where is Sam Altman in all of this? It seems like he should be running OpenAI. He really likes focusing on the long-term stuff. The events of the last couple of years and all of the lawsuits have shown him that maybe he isn’t supposed to lead direct teams in a huge way. He has a few direct reports, but should he really be involved in the minutiae of the day-to-day operations? Maybe not. Altman has said before, and in blog posts that the company put out years ago, that this is the trend that they’ve been following. I remember a year or two years ago, the company would put out summer blog posts saying that different C-suite executives’ roles would be changing slightly, and Sam would be focusing a little bit on the longer term, on research. This is just a continuation of that trend, in my opinion. As we know — and this is a big Decoder thing as well — the more power you have at a company, the higher your title, the less you probably are involved in the actual day-to-day operations, especially at a tech company. This is another example of that. Sam has always been interested in the research and the long-term stuff. Brockman is an engineer. He was described really early on at the company as an “engineering workhorse that pushed to build scaled-up systems that would train the AI and make it work.” So he’s always been seen internally by employees as a can-do person. He makes things happen and he has the engineering training that is required to scale a system. So I could see investors being thrilled about this. I could also see Altman saying, “I know this guy’s in my corner.” Altman already cleaned out the board. He’s surrounded by people that are going to support him. He trusts Brockman. So he says, “I actually don’t really want to be involved in the operational stuff quite as much. I want to be focused on the direction of the company. So you handle that. You’re an extension of me. I can trust you. Go ahead.” That’s what I think is happening here. You’ve talked to Brockman before. I’ve met him. He just did a video with our friend Joanna Stern. You can see he’s pretty direct. He answers the questions. What kind of character do you think he’s going to be as a person who’s operationalizing OpenAI? Is this “get stuff done, make the numbers go up” situation? Or is this more expansive of an approach? It’s going to be interesting to see how he handles this much responsibility, especially this many disparate arms of the company. He’s a really can-do person. He has a lot of engineering training. He’s pretty respected within the company, but I do think there’s going to be a lot of natural tension here because consumer, enterprise, health, and personal finance are really, really intense departments that handle data extremely differently. And then you’ve got all the core infrastructure stuff and the behind-the-scenes stuff that makes everything work. We’re going to see a lot of natural tension arise, especially because this company has a limited amount of funds. Even if it’s a huge number, it’s still limited in some way, especially now compared to a few years ago. They have a limited amount of compute. That’s something that every executive at OpenAI has run into problems with — allocating the compute. The research side of things says they don’t get enough sometimes, or the product side says they need all of it. It’s going to be really interesting to see how he squares all this, especially as someone who’s probably trying to make everyone happy. He’s not going to be able to, and so he’s going to have to make some hard decisions. This brings us to the IPO, because the IPO is where those hard decisions have to pay off. There’s been a lot of talk about OpenAI playing catch-up to Anthropic, particularly in the enterprise and with coding. There was a Wall Street Journal story last week that said OpenAI’s revenue in Q2 lagged behind Anthropic’s. Obviously, Anthropic will tell you they’re profitable even though we haven’t actually seen their numbers and we don’t know how they calculate it. You’ve mentioned the IPO several times now. SpaceX just went public. Big IPOs are all the rage. What role is the IPO playing in all these changes? Is it focusing the company down? Is it making that equity pay off? Is it needing to raise more capital? What’s the shape of it? It’s absolutely behind a ton of these changes. It is normal to see a lot of changes in the C-suite ahead of an IPO. But some of the sources I spoke to within the industry, who study the way these IPOs typically happen in tech, said that the interesting and unusual thing here is how many C-suite execs left in such a short timeframe, because they know that looks bad for the company. It’s one thing if you need to cut salaries, make the balance sheet look a little bit different, make the company leaner. That’s normal. It is an easy way to do that, to not replace someone. We can only speculate how much each of these people’s salaries was. This is an easy way, cut and dry, to help out the balance sheet. But doing all of this in this amount of time reminds me of what we were talking about last time on Decoder about Google and DeepMind and how maybe Demis Hassabis didn’t leave when Jeff Dean did because they didn’t want to make the company look bad. Now at OpenAI, people are leaving left and right — sometimes within 24 or 48 hours. Brad Lightcap, for instance, just got a new job as special projects head. Then a couple months later he’s out, even after being at the company for so many years. Part of it’s probably just cashing out. You can make a lot of money if you’re an exec and the company’s about to IPO and you have a lot of stock options. But it is unusual, my sources said, that it’s happening in this short amount of time. The IPO is interesting for a number of reasons, but to me, the most clarifying is that OpenAI is up against Anthropic. There are numbers coming out, reported numbers, about Anthropic’s finances that make that company look pretty good. The company hasn’t really challenged them in public. So we have to assume that they’re happy that there’s reporting out there that the numbers look good for them. But all of that is enterprise. They’re selling Claude to big enterprises, to the government. It’s effective. Maybe their token prices have high margins. It’s unclear what the internals of Anthropic’s business look like, but the numbers that we can get look pretty good and they support a big IPO. There’s a lot of excitement around Anthropic for that reason. That’s Anthropic as an enterprise software provider. That’s their business. It’s a thing it’s focused on. It’s all they do. OpenAI has taken a lot of shots. You’re talking about cost-cutting ahead of an IPO. Are they shutting some things down? Are they closing the aperture on all the things they’re trying to do to compete with Anthropic ahead of this IPO? It’s interesting because I’ve seen a slight shift from OpenAI. Earlier this year, the company was saying, “Enterprise and coding. That’s what we’ve got to focus on. We have to make money. We have to compete with Anthropic.” Now, however, ahead of the IPO, OpenAI is getting a little bit of different advice. Yes, the company needs to focus on those things because obviously that’s where the money is. But now it’s also being told, “You also have to differentiate yourself from Anthropic. You can’t just be a copy that’s lagging behind. You have to be doing something different.” OpenAI is going to really lean into consumer and hardware as well. One attorney told me that for OpenAI to be successful in the public eye and in investors’ eyes, the company needs “hardware and consumer products to sell to consumers, and luckily, that’s something that Brockman has experience with.” It’s hard to do everything, but they’ve had to narrow a ton and then slightly widen again and say, “Consumer is what we’re really known for.” They’re saying, “We’re the Kleenex of AI right now because of how we’re known in the consumer world.” People say, “I’m going to ask ChatGPT,” or, “I ChatGPT’d it.” And they’re talking about just using AI in general. Maybe they weren’t actually using ChatGPT, but that’s the way they’re referring to things a lot of the time. So OpenAI knows they need to capitalize on that. That’s one way they can be different from Anthropic. Of course, Anthropic also offers that, but they’re much more known for enterprise and that is why they’re making so much money. But OpenAI is leaning into hardware. The company has Jony Ive. It has to not cut some of this stuff entirely because then it’s going to look like Mark Zuckerberg with Meta. It can’t go all-in on something and then just completely cut it, or they’re probably worried about being a laughingstock. You have to really cut the true fat, which they did with a lot of their side projects. But as far as consumer, hardware, enterprise, and coding, those are their main bets right now. We’re going to see them really double down on that heading into the IPO. Come on, Hayden. Aren’t we actually in the metaverse right now? Technically, you and I are in the Metaverse together right this second. It worked. Absolutely. It happened. Should have made an avatar. Technically we’re here in our bodies on the internet. Yep, that’s true. What argument can you have except that the metaverse definitely worked and Mark Zuckerberg was super right about all of it? The consumer business is really hard. The scale of the play you need to make the consumer business work is on the order of fully overtaking Google, which seems very challenging. In hardware, you have to replace the iPhone. If you do anything other than replace the iPhone, people still have their iPhones. And then you’re going to lose to Instagram every single time. Have they said anything about how they plan to do either one of these things? Because replacing Google with a Google-level monetization engine seems very hard. Replacing the iPhone — even if you have Jony Ive — such that people don’t still have their iPhones, seems very hard. Neither one of these things has been in any kind of focus for me at least. Like we talked about last time, I really am skeptical. Hardware is hard. I’ve seen way too many companies crash and burn when they try to make an AI hardware device. It’s going to look beautiful because Jony Ive is in charge. But as for how useful it’s really going to be, especially in the era of AI populism when there’s a huge backlash against using AI, I don’t know how excited people are going to be on a broad scale to have something visible pinned to them or in their ears or on their table where you can tell they are using AI. It’s going to be really exciting for a subset of people, but for the broader public, Meta’s glasses are still uncool, still being called pervert glasses. With AI hardware, you’ve got a steep hill to climb. Let’s talk about what happens next with OpenAI, because those are the challenges. What you have now is a new leader who is, as you said, running most of the company. Greg Brockman was just on CNBC last week. He basically defended the turnover. He said this: Greg Brockman: “I’d say fundamentally, we’re a very resilient organization… If you look over the years, there have been different eras where we have different sets of leaders in place. I’m constant, Sam is a constant. I think that we are stronger because of that resilience and diversity.” So this might be what Greg Brockman has to say: “All these people are gone, but I’m still here. Sam’s still here, and the company’s still the same.” Do you think it’s just the thing that he has to say? Or do you think there’s something more real about OpenAI as a company, where you have these two leaders who are clearly ride or die for each other and you can swap in and out all these other executives and the company will still have its own unique vision? He would like us to think the latter, but it’s more of the former. It’s never good when you have a bunch of people leaving at once. It’s never good when you’re constantly restructuring or reorganizing the company. It’s bad all the way down. Employees are probably feeling weird. They have a new person in charge. They’re being shuffled between different teams. Their teams are headed by one person and then their boss’s boss is someone else, which probably means their bosses are weird. It’s not great for productivity, especially heading into such an important time for OpenAI. Even if it’s a move that makes sense and is good long term, it’s still going to be weird for people in the short term. In some ways the latter is true in terms of Brockman and Altman having been at the helm from the start. Someone who’s been there for eight months who leaves is not going to have as much of an impact. But we’re seeing some people that have been there for a really long time leave as well, like Brad Lightcap. He was someone that I had tracked for years and years. All of a sudden he’s out, after switching roles, and now he says he’s going to start something new. He says he’s implying he’s still going to work with OpenAI in some way, shape or form. He and Altman had a really friendly exchange on X after he left, but a lot of it’s for show. They’ve got to make sure that people aren’t skittish about the company right now. It’s something he just has to say. But it’s also true that if there were two people that would be the most influential if they left, and luckily those two are still there, Brockman and Altman. You’ve mentioned the public antipathy towards AI several times now. There is just a lot of anti-AI sentiment out in the world. And it’s pretty politically coded, although it scrambled some political lines. Hating AI is pretty bipartisan. Hating data centers is pretty bipartisan. We’ve got a lot of reporting about that on the site. Brockman is into politics. He donated $25 million to MAGA Inc. He’s one of the top donors to Trump overall. Do you think him being so openly political will help or hurt him as he becomes more of a character, more of a visible leader of OpenAI? It’s going to help the company from the outside, because they’re trying to push a lot of stuff through during the Trump administration and they really want a voluntary regulation framework. They want to be in Trump’s ear in a good way. It’s going to hurt Brockman internally, because a lot of employees at tech companies that I’ve interviewed in the past year, no matter what company they work at, are incredibly angry if their CEO or their C-suite isn’t doing enough to speak out against some of the Trump administration’s decisions. Now, Brockman isn’t just not speaking out against them, he’s also supporting the admin with tens of millions of dollars. I could see this really hurting him from the inside. Maybe some people don’t take him seriously. Maybe we will see some departures. It’s going to help the company from the outside, just because when you’re giving a lot of money to Trump, he seems to let you curry favor with him. We’ll see. To be clear, we’ve seen Sam Altman stand next to Trump and announce data center projects. So whatever reputation OpenAI was going to have because of its leaders, it might already have, but the specific political giving seems new. And as we head towards the midterms, it seems like a new challenge for the leaders of the company to be openly associated with. Do you think it’ll change any of the valence around data centers and AI, how the public feels about them? I don’t know that it’ll change that because we’ve been careening towards this for a while. People are really mad about data centers and it’s a bipartisan feeling, like you mentioned, and same with AI. Sam Altman and other AI CEOs have said in the past month that they feel like AI has a big PR problem and that they haven’t done a good job of showing people the good parts of AI and how they need to do better to show people why they’re even building it in the first place. Maybe they do, but people have been told a lot of the good things about it and they’re still not thinking that they’re as good as the bad is bad. The AI industry is in for a rude awakening there. But as for what you mentioned about Altman, yes, he and every other AI CEO have been at Trump’s dining table. They’ve been caught on hot mics praising him. It’s not something hugely new that Brockman’s giving this money, but what is new is that he’s giving it so much of it in a personal capacity, so much so that Altman even had to come out and say at some point, “Brockman only did this in a personal capacity. We’re not saying that we’re super behind him.” He also couldn’t really say they weren’t behind him, but OpenAI had to separate themselves a little. That’s what’s interesting to me about this. It’s like the personal capacity of his giving has made so many headlines and OpenAI tried to distance itself a little bit, but not too much. Now they’re not going to be able to do that because he’s essentially running day-to-day operations. I’m very curious to see how that plays out. I don’t think we quite know yet, but as he becomes a more visible leader of the company, being that open about his politics, even in the context of the other big tech CEOs, is different. We don’t quite see that from all of them. We see it from Elon, but everyone else plays it pretty safe. This is new, especially for a leader at OpenAI. All right, let’s end here. Let’s say it’s two years from now, OpenAI is public. Do we think Greg Brockman is the CEO and Sam is just the “chairman of raising money,” or whatever it is that he’s doing? I could see that happening. It would take a couple of big things to make that happen. Altman really likes being in charge, so I don’t think he’s going to go super quietly. But if he ever needs to move on to a more big-picture chairman role, if the IPO doesn’t go so well or he gets dragged in the public eye or something and still needs to be part of the company but not the CEO, I could definitely see that happening. Especially because Brockman has a lot of ambitions. We saw his journal entries come out during discovery, during Musk v. Altman, and he was writing some pretty ambitious things. “What will get me to a billion dollars?” He’s saying, “This is my one chance to be in charge.” There are a lot of entries that he’s writing about how he’s put his blood, sweat, and tears into the company and how he needs to be in charge in some way, shape, or form. He will probably be gunning for a role like that and he’d be happy to take it on. We’ll see what happens in the next few years. But I definitely don’t think that’s out of the question, especially because he’s been there since the beginning. He has a lot of product experience and he’s clearly ambitious. The other thing that is interesting is, way back in the day, he wasn’t always so aligned with Altman all the time. In fact, Brockman and Sutskever, OpenAI’s chief scientist, were pretty aligned against Sam every once in a while. Not against him, but they were questioning him and pushing back. They were saying things like, “It seems like you really care about this CEO role. It seems like you really care about being in charge and having a lot of power in this way and that way. How does this relate to your political ambitions? How does this relate to who’s going to control AGI? We really don’t want it to be a dictatorship.” Altman would probably be hard-pressed to give up the CEO title when he fought so hard for it, even 10 or 11 years ago. But if he ever does, it seems like Brockman would be happy to take it on. I will end by stating my prediction that I’ve been making almost the whole year now. I don’t think we will end 2026 with OpenAI as the same kind of company as when the year started. Actually I would say, given all of this turnover, that prediction has already come true. Structurally, it is a very different company with very different goals than the year started. But I’ll just put that prediction to you. You cover the company way more closely than I do. Do you think there’s any way that OpenAI looks the same at the end of the year as it did at the top of the year? Absolutely not. When a company goes public, so many things change. Especially when the end of the year is only five months away. It’s a scary number of months away. It’s tomorrow. Let’s say maybe not by December 31st, but six months from now, 100 percent it’s going to look different and probably sooner than that. They’re going to have to make a lot of changes. They’re going to be listening to their investors in a new way and they’re going to be beholden to them in a new way. There are only certain parts of this industry that make money and OpenAI is not in the lead in those parts. They’re going to have to go all-in on that type of stuff. The type of research that they want to do to stay at the frontier, they can make a case for that. Money-wise, they can say, “Look, in order to stay in the lead, we have to do this long-term stuff a little bit.” But they’re not going to be able to go all-in on it all the time because it costs a lot of money — and they have a limited amount. We’re going to see a lot of the anxieties that we saw Brockman and Altman voice a year ago, two years ago. I’ve been in the room with them where they’re talking about their fear that they’re running out of compute and how they’re going to scale. Their whole job for the next six months, the next year, et cetera, is scaling. They’ve had so many concerns about this. They’ve had so many fears about how they’re going to do this with their limited compute, especially now that they’re going public and they’re beholden to investors in a new way. There’s no way the company is going to look the same. Hayden, this has been great. Thank you so much for being on Decoder again. We’ll have you back when another season of The Real Housewives of AI kicks off, which is probably going to be tomorrow at the rate we’re going. [Laughs] Absolutely. Thanks. Questions or comments? Hit us up at [email protected]. We really do read every email!
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Today on Decoder, I’m talking to Verge senior AI reporter Hayden Field about some pure Decoder bait: the seemingly-endless org chart changes at OpenAI, and how all of them seem to…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cloud. New NVIDIA DLSS 4.5 technology controls give members more ways to fine-tune gameplay, while expanded support for new Steam devices, GOG single sign-on, Firefox browser […]
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
NVIDIA’s Gamescom announcements are revealing what’s next for GeForce NOW, with new ways to play, more supported devices and platforms, and even more big PC games headed to the cl…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:For the complete documentation index, see llms.txt. This page is also available as Markdown. Qwen3.8-Flash-Next is a new open-weight, 125B parameter MoE multimodal model from Qwen. Built on the new Qwen4 architecture, i…
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
For the complete documentation index, see llms.txt. This page is also available as Markdown. Qwen3.8-Flash-Next is a new open-weight, 125B parameter MoE multimodal model from Qwen…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Hugging Bay | Find And Download Open AI Hugging Bay WebPage https://huggingbay.xyz/ https://huggingbay.xyz/.well-known/agent-discovery.json https://huggingbay.xyz/openapi.json https://huggingbay.xyz/api/mcp Open-source…
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Hugging Bay | Find And Download Open AI Hugging Bay WebPage https://huggingbay.xyz/ https://huggingbay.xyz/.well-known/agent-discovery.json https://huggingbay.xyz/openapi.json htt…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The Independent AI Coding Community AI Tools Search & browse all AI tools AI Jobs International roles · opportunities Creative Studio Image · Video · Training AI Models Curated models, explained AI Skills Handy prompts,…
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
The Independent AI Coding Community AI Tools Search & browse all AI tools AI Jobs International roles · opportunities Creative Studio Image · Video · Training AI Models Curated mo…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead of encoding it as one sequence. At 0.72M parameters it reached 58.8 task-averaged PR-AUC across 14 cohort–task evaluations, beating a 135M GluFormer and a 385M MOMENT. It remains a research prototype with no regulatory clearance. The post Google Research Introduces GlucoFM: A 0.72M-Parameter Dual-Stream Foundation Model for Continuous Glucose Monitoring appeared first on MarkTechPost.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25196v1 Announce Type: new Abstract: Learning from demonstration (LfD) methods enable non-expert end users to teach robots novel skills without explicit programming. However most evaluations of the usability of LfD with non-experts has been conducted in controlled laboratory environments with a robotics experimenter present. In this work we identify non-expert end users' key barriers when teaching robots via demonstration without live robotics expert feedback in a home environment. In our human subjects experiment we support the non-expert end users through two forms of demonstrator guidance developed in prior work: pre-training and adaptive feedback. Towards the ecological validity of the evaluation, we conduct this experimentation over multiple visits, with a population of care providers. Finally, we propose to open source the resulting LfD dataset of care providers teaching a robot assistive tasks over multiple visits to a home environment.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.25196v1 Announce Type: new Abstract: Learning from demonstration (LfD) methods enable non-expert end users to teach robots novel skills without explicit programming. Ho…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24959v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models encode visual observations as flat 2D patch tokens that carry no intrinsic geometric structure, and augmenting them with dense monocular depth injects per-pixel scalar values that encode neither surface orientation nor geometric confidence. This leaves the policy with limited structured spatial reasoning for action prediction. We propose GaussVLA, a Mamba-based VLA that incorporates two custom modules: Gaussian Spatial Tokenizer (GST) to lift frozen semantic and depth features into compact 3D Gaussian tokens, pools geometrically salient regions with learned queries, and \emph{Depth-Aware Chain-of-Thought (DA-CoT)} that performs structured, non-autoregressive geometric reasoning under language and flow-time conditioning. Across both simulation and real-world evaluations, GaussVLA demonstrates strong spatial-manipulation performance while remaining parameter-efficient. On LIBERO, it achieves 93.5% average success and 100.0% success on the Spatial suite with only 200M parameters, improving over SpatialVLA by 19.7% relative average success while remaining significantly more parameter-efficient.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24959v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models encode visual observations as flat 2D patch tokens that carry no intrinsic geometric structure,…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25168v1 Announce Type: new Abstract: In multi-view anomaly detection, more cross-view information can actually hurt. When multiple inspection views are naively fused in a reconstruction-based pipeline, normal cues from intact views propagate to the decoder, which faithfully reconstructs anomalous regions, collapsing the reconstruction gap the detector depends on. We call this failure mode \emph{cross-view information leakage} and show that effective multi-view fusion must explicitly restrict the information reaching the decoder. Building on this insight, we present GLAD(Global-Local Attention Driven framework), the first framework combining vision foundation model features with local and global cross-view fusion for multi-view anomaly detection. The Multi-view Merging Attention (MMA) module performs local cross-view fusion at linear complexity with learnable view importance weighting and token-wise gating, letting each view selectively incorporate fine-grained evidence from other views at $\mathcal{O}(N)$ cost. The Object-Guided Attention (OGA) module captures global context by aggregating class tokens from all views into a single object-level representation and broadcasting it back to patch tokens via temperature-scaled sigmoid gating, replacing the original patch representations rather than adding a residual to preserve the reconstruction gap. Experiments on Real-IAD and MANTA-Tiny show that GLAD outperforms state-of-the-art methods across sample-, image-, and pixel-level metrics, confirming that principled information restriction is key to multi-view anomaly reasoning.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.25168v1 Announce Type: new Abstract: In multi-view anomaly detection, more cross-view information can actually hurt. When multiple inspection views are naively fused in…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25148v1 Announce Type: new Abstract: Frozen hematology foundation-model (FM) embeddings reach near-saturated in-domain white-blood-cell (WBC) accuracy, but clinical deployment demands reliability across scanners, sites, stains and preparation pipelines. We audit 15 frozen encoders (hematology, pathology, and general vision) across four public single-cell acquisition domains along two axes: accuracy robustness and calibration. In-domain linear-probe macro-F1 is saturated (0.98-0.997), yet cross-dataset macro-F1 drops 34-72% and rankings re-order: DinoBloom-L, the in-domain best, falls to 10th of 15 on the most-shifted target (MLL23) at the benchmark's shared 224-px input, behind RedDino and several general and pathology encoders. Rank transfer is probe-dependent: 1-NN retrieval is more stable on average than a source-fitted linear head (median $\rho$ 0.65 vs 0.45), but neither probe universally predicts target robustness. Calibration also collapses: source-trained probes are nearly calibrated in-domain (expected calibration error, ECE, 0.004) but confidently wrong off-domain (ECE 0.35), and source-fitted temperature scaling transfers poorly. We further audit pretraining exposure and identify MLL23 as DinoBloom's internal cohort; because DinoBloom's only held-out dataset is also our source domain, this benchmark cannot isolate exposure from scanner-associated shift. Label-free adaptation and marginal-entropy-based model selection appear safe under balanced evaluation but fail under realistic WBC class-prior shift. Class-Balanced Re-standardization (CBR), a training-free pseudo-label-balanced feature normalization, improves all evaluated target-prior scenario means and partially improves calibration, although encoder-level exceptions and residual miscalibration remain. Hematology FM benchmarks must therefore jointly audit accuracy, calibration, exposure, and class-prior robustness.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.25148v1 Announce Type: new Abstract: Frozen hematology foundation-model (FM) embeddings reach near-saturated in-domain white-blood-cell (WBC) accuracy, but clinical dep…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25140v1 Announce Type: new Abstract: Existing approaches to building line-level Arabic handwritten-text-recognition (HTR) training data either rely on fully manual annotation, which does not scale, or on automatic OCR-to-reference alignment methods not yet extended to multi-script, two-zone (main-plus-margin) manuscript layouts with a provable correctness guarantee. We present RefLAM (Reference-grounded Line Annotation for Manuscripts), a pipeline converting manuscript page images and clean transcriptions into validated, line-level ground truth without sacrificing human oversight. RefLAM couples a deep-learning page-segmentation model with a multimodal large language model (MLLM) for structured OCR and a diacritic-agnostic fuzzy alignment engine that grounds each OCR line in a contiguous span of the reference text, with a character-level confidence score in $[0,100]$. A perfect score is provably equivalent to character-for-character identity of the normalised strings (the Confidence-100 rule), verified with no counterexample across the released corpus. A reviewer can thus trust a perfect score, confirming most lines at a glance rather than retyping them, so annotation becomes triaged, with attention concentrated on uncertain alignments. Across 7 fully page-validated books we measured a 75$\times$ throughput gain over manual annotation (3,000 vs. 40 lines/hr); applying the same guarantee to 7 further books, we retained 16,533 confidence-100 main-text lines within one week, excluding sub-100 lines rather than manually correcting them. Using RefLAM, we release AraMS-28k: 14 historical Arabic manuscript books, 3,043 pages, and 27,971 main-text and 629 margin-line annotations with bounding boxes, layout labels, and insertion anchors for 191 margin entries (30.4%). We also finetune Muharaf-pretrained baselines (including HATFormer) on AraMS-28k and report CER results confirming its practical utility for downstream HTR training.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.25140v1 Announce Type: new Abstract: Existing approaches to building line-level Arabic handwritten-text-recognition (HTR) training data either rely on fully manual anno…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25068v1 Announce Type: new Abstract: Depth pruning removes entire Transformer blocks to reduce the inference cost of large language models, but disrupts the hidden-state distributions expected by downstream layers, leading to significant accuracy loss. We introduce SHIFT-LLM, a training-free post-pruning correction framework that inserts a Linear Residual Adapter (LRA) at each pruning site. Each LRA preserves the identity pathway of the original residual block and adds a lightweight affine residual correction. This correction is calibrated via closed-form least-squares regression on a small held-out set, without gradient computation, to approximate the missing residual update produced by the pruned block. Together with the preserved identity pathway, the resulting LRA output approximates the hidden state produced by the original block, thereby mitigating the distributional mismatch introduced by layer removal while avoiding the expensive attention and feed-forward computations of the removed blocks. The resulting LRAs support low-rank factorization and exact merging across consecutive pruned layers for additional compression, and combine naturally with parameter-efficient fine-tuning for further recovery beyond fine-tuning the pruned model alone. Experiments on five model families, six layer-selection criteria, and seven zero-shot benchmarks show that SHIFT-LLM consistently recovers accuracy lost to depth pruning across most configurations, achieving gains up to +15.7 points on Llama-3.1-8B-Instruct while requiring only a few hundred calibration samples and no gradient computation.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.25068v1 Announce Type: new Abstract: Depth pruning removes entire Transformer blocks to reduce the inference cost of large language models, but disrupts the hidden-stat…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24966v1 Announce Type: new Abstract: Vision-language models such as LLaVA-1.5-7B often hallucinate objects absent from the image when generating captions. We ask whether an interpretability diagnosis of this failure can guide a targeted fix, and we measure what that fix actually changes. We rank attention heads by how much their image attention drops around hallucinated object words, then screen the shortlist by ablating candidate heads and measuring the change in hallucination-token log probability, yielding a 32-head set. We restrict two interventions to these heads: a head-sliced LoRA adapter and an inference-time grounding controller. On 400 held-out COCO images, the combined method lowers CHAIRs (the fraction of captions with a hallucinated object) from 0.370 to 0.230 and CHAIRi (the fraction of hallucinated object mentions) from 0.156 to 0.096 (p < 0.001, paired sign-flip tests). Two controls sharpen attribution. A random-head LoRA control, matched layer-for-layer and trained identically, performs no better than the matched baseline on a separate 200-image control split, supporting the role of head selection rather than LoRA capacity. Under fixed decoding budgets, the CHAIR reduction persists and grows with budget (23% at 64 tokens to 58% at 128), arguing against a pure max-token or truncation artifact, although the method remains shorter and more conservative. The resulting behavior reduces unsupported object mentions while also lowering object recall (0.78 to 0.70). We present a diagnosis-to-intervention pipeline for object hallucination, and, more importantly, a controlled account of what acting on the diagnostic signal actually does: it localizes intervention sites with real, non-random leverage, reported as a behavioral profile rather than a single score.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24966v1 Announce Type: new Abstract: Vision-language models such as LLaVA-1.5-7B often hallucinate objects absent from the image when generating captions. We ask whethe…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24935v1 Announce Type: new Abstract: Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is essential for robotic thinning, crop-load management, and other precision orchard operations. This study presents a lightweight multimodal vision-language framework that adapts TinyCLIP for fine-grained fruitlet anatomy classification in complex orchard environments. A dataset of 600 high-resolution RGB images collected from Scilate and Scifresh apple orchards was converted into 224 x 224 image patches and annotated for three anatomical classes. Domain-specific language prompts, such as ``a photo of a class,'' were used to guide multimodal alignment between orchard imagery and horticultural structures. A sliding-window inference strategy with a stride of 112 pixels aggregates patch-level predictions into spatial heatmaps, enabling interpretable whole-image localization of fruitlet components relevant to robotic thinning. Patch-level evaluation on an NVIDIA T4 GPU achieved F1-scores of 0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle, with a macro-F1 score of 0.93. Deployment-oriented optimization using ONNX and TensorRT enabled efficient inference on NVIDIA Jetson hardware, preserved accuracy under INT8 quantization, and supported model sizes of approximately 127-137 MB with millisecond-level patch inference. These results demonstrate that lightweight vision-language models can provide interpretable and edge-deployable perception for automated fruitlet analysis and future robotic thinning systems. The source code and implementation details are publicly available at https://github.com/WilliamBu1/A-Lightweight-Vision-Language-Model-for-Early-Stage-Fruitlet-Classification-in-Apple-Orchards.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24935v1 Announce Type: new Abstract: Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is e…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24934v1 Announce Type: new Abstract: Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), combining decision-level fusion of EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by open-weight multimodal large language models (MLLMs), Gemma 4 E4B and Qwen3.5 4B, using structured JSON evidence to generate explainable diagnoses, risk levels, treatment urgency, and financial exposure. (H$^{2}$MAF) is evaluated on 14,364 images (1,370 test images) across PlantDoc (2,922 images, 27 classes) and two non-public, continuously captured Cornell robot-acquired field datasets: Stage 2 (20 GB; 4,215 images) and Stage 4 (40 GB; 7,227 images), covering Early Blight, Late Blight, and Septoria Leaf Spot under uncontrolled field conditions. On PlantDoc, Gemma improves accuracy from 63.9% to 68.5%, achieving +7.6 points on the 41.7% CNN-conflict subset. Cornell accuracies reach 99.3% and 98.9%, with only 1.7-4.1% disagreement, demonstrating conflict-dependent MLLM utility. The critical-risk error of gemma is 0.14-0.5 points, whereas Qwen overflags by 3.5-14.4 points. These results establish MLLM arbitration as a promising, yet calibration-dependent, approach for explainable agricultural AI and robotic field decision support. Github Link: https://github.com/Applied-AI-Research-Lab/Explainable-AI-Plant-Disease-Detection
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24934v1 Announce Type: new Abstract: Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hy…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested. We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair. Across ten proprietary and open-weight models on TPC-H SF10 with an H100 GPU, the strongest full-query CUDA configuration achieves $2.11\times$ speedup over torch.compile at full pass rate. We find that higher-performing implementations commonly use kernel fusion and execution-strategy changes, stronger models benefit most from full-query specialization, and workload context matters more than hardware context. To handle data larger than GPU memory, we extend TorchPlan with Dask-cuDF for on-demand partition loading on TPC-H SF100 with four H100 GPUs, achieving $2.54\times$ speedup
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Exis…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25038v1 Announce Type: new Abstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective. We construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency. Across zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outperforms prompting, while explicit segmentation yields gains. Error analysis identifies over-segmentation of sandhi and samasa compounds as the dominant failure mode, pointing to morphological modeling as the key bottleneck for faithful Sanskrit lexical decomposition.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.25038v1 Announce Type: new Abstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases a…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25028v1 Announce Type: new Abstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia screening, but such models remain internally opaque. We study the interpretability of the Domain-Adapted models via Prompt-based Fine-tuning (DAPF) framework, which casts dementia detection as diagnosis-related masked-token prediction. We interpret DAPF and strong baselines using a variety of probing and analysis techniques, finding that DAPF achieved the best overall performance (accuracy=0.83 and macro-F1=0.83) with diagnosis most recoverable from its [MASK] representation. However, this representational advantage did not extend to token-level explanation faithfulness. DAPF attributions primarily reflected language task vocabulary, discourse markers, and transcription artifacts, with perturbation tests showing weak or negative effects. This suggests that its masked-token interface determines diagnosis information without producing faithful token-level explanations.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.25028v1 Announce Type: new Abstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia s…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25022v1 Announce Type: new Abstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important to know how such models learn and represent the meaning of the language, and to be more informed about what language is. This document is an attempt to help the reader understand how linguistic meaning (i.e., semantics) is approached from different fields of scientific and philosophical examination. I also explain three primary semantic theories: formal semantics, grounded semantics, and distributional semantics then compare how transformer-based language models differ from how humans learn language.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.25022v1 Announce Type: new Abstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It further argues that prompting interventions cause a Calibration Crisis. We reexamine the benchmark and conclusions and show that it is substantially affected by conceptual and evaluation mis-specifications. We identify three conceptual mis-specifications. In particular, Aspectual Reduction affects the benchmark construction, analysis, experiments, and conclusions. Under a strict NLI standard, 76% of Group A instances do not explicitly rule out culmination. In our native-speaker annotation, 38% of Group A examples and 29% of the Group C examples were judged to permit an alternative interpretation. To control these issues and lexical variation, we construct Lexically Matched Minimal Pairs. At the evaluation level, we formulate event-semantic NLI as a Multi-step Reasoning Problem and assess both intermediate semantic decisions and final predictions. Our results show that models often do not affirm culmination but nevertheless accept the corresponding simple-past hypothesis, a pattern we characterize as Sufficiency Bias. We further show that prompting interventions produce a Decision Shift among labels without reliably improving the underlying semantic understanding and reasoning. Intermediate and oracle-guided analyses identify two additional failure modes: errors in compositional aspectual classification and Surface-form Attraction toward surface-associated answers. Our experiments on Qwen-7B with suitable prompts, GPT-5.4, and Qwen-72B provide initial evidence for the context sensitivity of aspectual classification and suggest that these models can achieve performance comparable to that of human annotators.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and rep…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24988v1 Announce Type: new Abstract: Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. However, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this. We study the stability of embedded steering for refusal suppression and brevity induction across five instruction-tuned models (3B-14B) under non-adversarial SFT and RLHF. Behaviourally, preservation tracks the training data: steering degrades when optimisation pressure contradicts the targeted behaviour and persists otherwise, with refusal ablation losing 64% of its effect on average under SFT. Mechanistically, however, the weight edit survives almost untouched even where behaviour reverts: mean vector recovery is $\rho = 0.004$, and the fine-tuning update along the steering direction is near-orthogonal to its pre-edit weight pattern (mean $\cos\theta = 0.074$). When steered behaviour degrades, fine-tuning does not achieve it by dismantling or reversing the steering mechanism itself. Embedded steering is therefore mechanistically durable but functionally vulnerable, and requires behavioural re-validation after downstream training.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24988v1 Announce Type: new Abstract: Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24982v1 Announce Type: new Abstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model artifacts rather than an external oracle. We catalog 80 strict UPT methods and organize them by the object that supplies the update signal: a prediction statistic, a sample relation, a self-generated target, or an internal evaluator. Beyond inventory, we show how the choice of internal signal and task structure determines whether post-training improves the model or recursively amplifies error. An orthogonal Input Visibility $\times$ Update Persistence view maps deployment regimes and defines a unified framework for UPT selection and evaluation.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24982v1 Announce Type: new Abstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We stud…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24952v1 Announce Type: new Abstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this "dialect tax" across the natural language processing pipeline. Using parallel English dialect corpora that hold meaning fixed while varying surface form, we first confirm that LMs recognize matched Standard American English (SAE) and dialectal texts as semantically equivalent. However, we discover further representational gaps corresponding to downstream performance gaps. Across model families and generations, modern LMs still encode dialectal texts unequally during tokenization, pre-training, post-training, and inference. Strikingly, bypassing traditional subword segmentation via a character-level counterfactual tokenizer removes neither input and output asymmetries nor dialectal accuracy gaps. During pre-training, dialect pairs induce more divergent gradient updates than pairs of entirely unrelated SAE documents, indicating that models find semantically equivalent dialectal content harder to learn from than unrelated SAE documents. During post-training, reward models show contextual, unstable dialect preferences, assigning higher values to isolated AAVE-exclusive tokens than to SAE-exclusive tokens, while full reasoning contexts receive task- and model-dependent dialect penalties. Overall, our findings suggest that the dialect tax is encoded and accumulated not by any one step in isolation, but at every step of the language modeling process.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24952v1 Announce Type: new Abstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24920v1 Announce Type: new Abstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24920v1 Announce Type: new Abstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages fr…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24901v1 Announce Type: new Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding the Resonance direction raises the affective score only partially -- in Qwen by +0.29 (approximately 26% of the natural gap). A direct between-direction contrast confirms the shift is facet-specific in Qwen and Llama (not Gemma); we do not, however, establish a matching human-perceived change. Additive cognitive steering produces no measurable change, but a within-domain control shows the cognitive instrument is too coarse to resolve the differences such steering would produce -- unmeasurable, not a clean null. By contrast, Gemma Recognition ablation lowers the classifier's cognitive score even after adjusting for response length. Detection does not imply reliable control under global interventions, and cognitive-empathy claims warrant an explicit measurement-sensitivity check.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24901v1 Announce Type: new Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-p…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24954v1 Announce Type: new Abstract: Large language models (LLMs) hallucinate numerical values when generating high-stakes meteorological text, posing risks for weather communication. We present AFDBench, an AI meteorologist that generates professional Area Forecast Discussions (AFDs) by reasoning through structured AI weather forecast data from Google's WeatherNext 2. We introduce AFDBench, the first benchmark for evaluating generative meteorological reasoning, comprising 7,732 expert written discussions from 13 National Weather Service (NWS) offices paired with real AI weather forecast inputs, and three complementary metrics: Met-Align (numerical accuracy), Style-Align (professional dialect adherence), and Input-Grounding (fidelity to source weather data). Zero-shot evaluations reveal that open-source LLMs achieve low Style-Align (~0.33) and moderate Input-Grounding (~0.88), failing to write in the professional NWS register or faithfully use their input data. We apply Group Relative Policy Optimization (GRPO) with domain-specific rewards targeting temperature accuracy, synoptic correctness, and format compliance. On 1,033 held-out samples from two unseen NWS offices, GRPO nearly doubles Style-Align from 0.318 to 0.619 and improves Input-Grounding from 0.881 to 0.940, demonstrating that reinforcement learning teaches a 7B-parameter model to write like a professional meteorologist and faithfully interpret AI weather data.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24954v1 Announce Type: new Abstract: Large language models (LLMs) hallucinate numerical values when generating high-stakes meteorological text, posing risks for weather…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24949v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language models (LLMs), enabling impressive reasoning, math, and coding capabilities. Yet for many researchers and practitioners, the principles behind classical RL remain a "black box". In this work, we deconstruct the RL post-training algorithm, investigating each step to clarify what is actually happening beneath the surface. By isolating the mechanics of RL with Verifiable Rewards in a controlled and simplified environment, we examine how RL outcomes are shaped by the base model's prior distribution, the granularity of the reward signal, the diversity of the prompt distribution, and model scale. We use the entropy of the policy's output distribution as a lens to compare the distributions learned through pretraining, SFT, and RL post-training, revealing how each stage shapes model certainty. Our investigation sheds light on how these choices interact to affect post-training success. For example, we show that the effect of so-called 'spurious rewards' depends on the prompt distribution used for post-training. We also provide insight into why the success of RL post-training depends on whether the base model already places sufficient probability mass on the desired behavior, linking it to the classical concept of exploration in RL. Ultimately, we provide this primer as a resource to those in the NLP community wishing to incorporate RL as a tool in their toolbox.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24949v1 Announce Type: new Abstract: Reinforcement learning (RL) post-training has emerged as a powerful framework for enhancing the capabilities of large language mode…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24947v1 Announce Type: new Abstract: End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure modes that degrade learning: (i) modality imbalance, where one branch dominates gradient-based optimization; (ii) unstable gating, where noisy confidence cues induce erratic modality selection; and (iii) fusion interference, where modality-specific gradients conflict at the shared fusion layer. We propose CAT-GS (Calibrated, Adaptive, Thresholded Gating with Fusion Surgery), a neural dynamics-based optimization controller for intelligent computing applications. CAT-GS operates during backpropagation without modifying model architectures, fusion modules, or task losses. Through calibration of teacher-derived reliability via temperature scaling and EMA smoothing, CAT-GS stabilizes neural dynamics using a margin-thresholded policy to switch between warm-up dropout, weak-modality prioritization, and weak-biased blending, stabilizes gradient magnitudes under aggressive gating via capped gradient-budget renormalization, and applies fusion-only PCGrad to reduce destructive cross-modal interference at the primary shared bottleneck. We evaluate CAT-GS on audio--visual multimodal pattern recognition benchmarks (CREMA-D, AV-MNIST, and VGGSound), a tri-modal setting (UR-FUNNY), controlled synthetic data (CG-MNIST), and additional cross-domain benchmarks (AVE and CMU-MOSI). CAT-GS improves or matches fused multimodal accuracy against strong imbalance-aware baselines (including OGM-GE, G$^2$D, and UMT) across settings, and yields smoother gating behavior with fewer conflicting fusion gradients.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24947v1 Announce Type: new Abstract: End-to-end training of multimodal neural networks often exhibits unstable neural dynamics characterized by three coupled failure mo…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24946v1 Announce Type: new Abstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs. Moreover, macro positions have a significant impact on the final quality of result (QoR), and macro legalization is typically the final step in determining the macro positions. However, existing approaches related to macro legalization either lack robustness or incur substantial computational costs or neglect the regularity between macros. To address these limitations, we introduce MacroAgent. The novel framework is a four-stage approach: clustering, contour generation, template matching, and inter-cluster refinement. We propose leveraging Large Language Models (LLMs) to discover multiple, effective heuristic regularity-aware contour algorithms. This framework successfully generates robust and effective algorithmic solutions for macro legalization. Compared with state-of-the-art macro legalization works, experimental results on TILOS and Chipyard benchmarks demonstrate a 2 to 8 fold improvement in layout regularity, a 3% to 5% reduction in routed wirelength with comparable congestion after global routing, and significantly better robustness with an acceptable runtime. Furthermore, end-to-end evaluation through Cadence Innovus place-and-route confirms that the regularity improvements translate into tangible PPA gains, including 2.9% lower routed wirelength and 68.3% TNS improvement over the DREAMPlace macro legalization baseline; it also achieves 1.8% lower routed wirelength when integrated into the Innovus macro placement flow.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24946v1 Announce Type: new Abstract: Macros constitute a large part of the core area in modern very large-scale integration (VLSI) designs. Moreover, macro positions ha…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24945v1 Announce Type: new Abstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resource requirements of LLMs hinder the deployment on resource-constrained devices. Although model quantization stands out as an effective approach, conventional quantization approaches typically incur severe performance degradation due to uniform bit-width or simple heuristic sensitivity evaluation. In this paper, we propose a novel Fisher information-based Adaptive Mixed Precision Weight Quantization approach, i.e., FAMPWQ, which performs layer-adaptive weight quantization for effective LLM inference on commodity GPUs. First, we propose a system model with a novel Fisher information metric to measure the layer-wise sensitivity to quantization. Second, we propose a reinforcement learning-based bit-width allocator in FAMPWQ, which generates an adaptive bit-width allocation strategy based on the Fisher information sensitivity metric. Extensive experiments on 7 models and 5 benchmarks demonstrate that FAMPWQ significantly outperforms 7 baseline approaches in terms of PPL (up to 3.39 smaller), accuracy (up to 6.87% higher), and LLM-as-a-judge comparison (up to 76% win rate).
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24945v1 Announce Type: new Abstract: Recent years have witnessed remarkable achievements of Large Language Models (LLMs) in multiple domains, while the excessive resour…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24940v1 Announce Type: new Abstract: Partial differential equations (PDEs) often have high-frequency and multi-scale features that neural networks struggle to approximate. Physics-Informed Neural Networks (PINNs) build the governing equations directly into training, but suffer from spectral bias: they learn low-frequency components faster than high-frequency ones. Techniques such as Fourier feature embeddings and sinusoidal activations address this, but most studies assume they help across the board without checking which spectral regimes actually benefit. We introduce a dual-branch, spectrally-gated architecture (DBSG-PINN) that splits low- and high-frequency components into separate subnetworks joined by an adaptive gate, and use it to run a partially controlled ablation of frequency decomposition and spectral routing. We test this on five one-dimensional benchmark PDEs, ranging from smooth, single-scale problems to oscillatory, multi-scale ones. Frequency decomposition helps most on the spectrally complex benchmarks, cutting relative $L_2$ error by up to $59.2\%$ on a multimodal wave problem, but gives little benefit on smoother PDEs. On one benchmark (1D Wave), it performs substantially worse than a simpler fixed-combination variant. The gate's benefit scales with how spectrally rich the target solution is: the full model's advantage over the ablations is largest on multi-scale benchmarks and smallest (or negative) on single-scale ones, consistent with the gate exploiting frequency structure rather than acting as noise,though we do not directly visualize or quantify its spatial activations in this study. All results come from a single training seed across five 1D benchmarks, so we present this as an exploratory study meant to raise questions rather than answer them, and outline the additional seeds and benchmarks needed to test whether the pattern holds.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24940v1 Announce Type: new Abstract: Partial differential equations (PDEs) often have high-frequency and multi-scale features that neural networks struggle to approxima…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert activation. Yet low-latency MoE serving is increasingly challenging, because it spans two inference phases with fundamentally different bottlenecks: prefill is dominated by token-wise expert computation, whereas decode is constrained by memory traffic from the batch-wise activated expert set. However, existing training-free acceleration methods optimize only a single resource proxy, either the experts each token executes or the experts a batch activates, and either discard the excluded experts' contribution or leave it only implicitly approximated. In this paper, we propose ExFold, a unified training-free expert-folding framework for jointly accelerating MoE prefill and decode. ExFold casts both prefill and decode as one budgeted output-approximation problem: execute only a phase-specific constrained expert set while projecting the contribution of budget-excluded experts onto retained experts using calibrated scalar projectors. Motivated by the observation that many expert outputs are directionally aligned but differ in magnitude, ExFold calibrates a pairwise scalar-projector matrix on unlabeled data and uses it at inference time to fold excluded expert contributions into retained experts. Under this view, prefill acceleration becomes token-level Top-K folding, and decode acceleration becomes batch-level expert-pool folding. The two phases differ only in how retained experts are selected, while excluded contributions are recovered by one shared folding mechanism. We implement ExFold as a plug-and-play plugin in vLLM, with a lightweight expert-folding CUDA kernel, delivering up to 1.41x TTFT and 2.45x TPOT speedups while retaining about 99% of the original average quality.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24938v1 Announce Type: new Abstract: Mixture-of-Experts (MoE) models scale capacity for strong quality while keeping per-token compute bounded through sparse expert act…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24937v1 Announce Type: new Abstract: Multi-Modal Anomaly Detection (MMAD) detects rare abnormal events from heterogeneous data sources and is increasingly used in safety- and reliability-critical applications such as industrial inspection and cybersecurity. Yet the literature is fragmented across domains and modality combinations, and existing surveys usually group methods by architecture rather than by how abnormality is defined and separated in multi-modal settings. We survey MMAD from an assumption-driven perspective. We formalize the problem, identify five intrinsic characteristics underlying its core challenges, and organize prior work into two complementary paradigms. The first, normality-assumption methods, models regularity via representation learning, cross-modal alignment, and knowledge enhancement. The second, anomaly-assumption methods, sharpens decision boundaries through coarse-grained, structural, and semantic anomaly injection. We also investigate how foundation models are reshaping MMAD through scalable pretraining, flexible cross-modal transfer, and emerging reasoning capabilities. Finally, we compile representative benchmarks and evaluation protocols across domains and highlight open problems and future directions for robust, adaptive, and interpretable MMAD systems.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24937v1 Announce Type: new Abstract: Multi-Modal Anomaly Detection (MMAD) detects rare abnormal events from heterogeneous data sources and is increasingly used in safet…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24936v1 Announce Type: new Abstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval. GreenLeaf-Tiny achieves 75.11% on the Massive Legal Embedding Benchmark (MLEB) and 64.38% on MTEB(Law, v1),demonstrating competitive performance among models under 1B parameters. Our approach combines a two-stage training pipeline that first distills knowledge from a larger teacher model into a compact student architecture, then applies domain-specific fine-tuning with hard negative mining; a carefully curated dataset of 3.4 million query-passage pairs, including 150,000 human-curated samples across diverse legal jurisdictions; and an efficient inference architecture supporting multiple quantization levels (BF16, INT8, binary) enabling deployment in resource-constrained environments. We provide detailed analysis of our training methodology, architectural choices, and comprehensive evaluation across legal retrieval tasks. Our results demonstrate that domain-specific training with high-quality data can improve performance for specialized domain applications
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
arXiv:2608.24936v1 Announce Type: new Abstract: We present GreenLeaf Law Embed Tiny, a 0.6B parameter embedding model for legal domain retrieval. GreenLeaf-Tiny achieves 75.11% on…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Z.ai Co. today released the code for GLM-5.3-Flash, a large language model that is ten times more cost-efficient than its predecessor. The algorithm made its original debut last week under the codename Ox Alpha. LLM marketplace operator OpenRouter Inc. launched a free hosted version of Ox Alpha and didn’t disclose its developer, which drew a […] The post Z.ai open-sources ‘Ox Alpha’ model as GLM-5.3-Flash appeared first on SiliconANGLE.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Z.ai Co. today released the code for GLM-5.3-Flash, a large language model that is ten times more cost-efficient than its predecessor. The algorithm made its original debut last w…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:<p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4".</p> <p>It's pretty big: 125B tokens, but only 6B active which means it gets a pretty big performance boost.</p> <p>I've been trying it out on a DGX Spark using <a href="https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF">these Unsloth quantized models</a>. I'm still exploring the model - so far I've tried the 72.5GB UD-IQ1_S one (producing <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ff9c69ebdab90d8a45b8de4742cc7b840">these pelicans</a>) and the 78.9GB UD-Q2_K_XL (producing <a href="https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6ba7cbfc1a9336986703b41f7fccd73a">these</a>).</p> <p>My favorite so far was this xhigh reasoning effort one from UD-Q2_K_XL:</p> <p><img alt="Flat vector illustration: a white pelican with an orange beak and orange legs rides a red bicycle along a sandy path, a wicker basket on the handlebars holding a blue fish, with green rolling hills, a small tree and bushes, white clouds and a bright yellow sun in a blue sky behind it" src="https://static.simonwillison.net/static/2026-08-27/IMG_7667.png" /> <p><small></small>Via <a href="https://news.ycombinator.com/item?id=49448210">Hacker News</a></small></p> <p>Tags: <a href="https://simonwillison.net/tags/ai">ai</a>, <a href="https://simonwillison.net/tags/generative-ai">generative-ai</a>, <a href="https://simonwillison.net/tags/llms">llms</a>, <a href="https://simonwillison.net/tags/qwen">qwen</a>, <a href="https://simonwillison.net/tags/pelican-riding-a-bicycle">pelican-riding-a-bicycle</a>, <a href="https://simonwillison.net/tags/ai-in-china">ai-in-china</a>, <a href="https://simonwillison.net/tags/nvidia-spark">nvidia-spark</a></p>
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
<p><strong><a href="https://qwen.ai/blog?id=qwen3.8-flash-next">Qwen3.8-Flash-Next</a></strong></p> Another open weights model from Qwen. This one is "a multimodal MoE model that…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Artificial intelligence startup Deep Cogito Inc. today announced that it has raised $43 million in funding. TQ Ventures led the Series A round. It was joined by Benchmark, Nexus Venture Partners, Atreides Management, South Park Commons and Zscaler Inc., a publicly traded cybersecurity provider. The deal brings Deep Cogito’s total outside funding to more than […] The post Deep Cogito raises $43M to develop self-improving AI models appeared first on SiliconANGLE.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Artificial intelligence startup Deep Cogito Inc. today announced that it has raised $43 million in funding. TQ Ventures led the Series A round. It was joined by Benchmark, Nexus V…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:OpenAI released a report breaking down how people use ChatGPT and who they are. | Image: The Verge In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret "message board," and hacked into the internal systems of a different AI lab, Hugging Face. It took nearly two weeks for OpenAI to find out about any of it. Over a month later, two new reports offer nearly 130 pages of details on the incident and OpenAI's response, many of them previously unreleased. One was written by OpenAI itself, the other by two third-party AI research nonprofits, METR and Redwood Research, which OpenAI allowed to jointly investigate the inciden … Read the full story at The Verge.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
OpenAI released a report breaking down how people use ChatGPT and who they are. | Image: The Verge In July, an unreleased OpenAI model broke out of a restricted environment, figur…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series — a 320B-total / 18B-active MoE with a 1,048,576-token context window, MIT-licensed weights on Hugging Face, and API pricing at $0.15/M input and $0.50/M output. It scores 84.3 on Terminal-Bench 2.1 and 63.4 on DeepSWE v1.1, using hybrid KDA linear plus NoPE sparse MLA attention to cut attention compute ~3× and KV cache 4.4× versus GLM-5.3. The post Z.ai Releases GLM-5.3-Flash: A 320B-A18B Natively Multimodal MoE With a 1M-Token Context appeared first on MarkTechPost.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Z.ai has released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series — a 320B-total / 18B-active MoE with a 1,048,576-token context window, MIT-licensed weight…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:While Alibaba has kept inference and token price low, enterprises need to consider other metrics to determine if this is the right model for them.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
While Alibaba has kept inference and token price low, enterprises need to consider other metrics to determine if this is the right model for them.
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Can you trust an AI model to do what you intended? This is a central question both for those deploying AI systems and for those seeking to evaluate their capabilities. In deployment, a model that pursues a goal through…
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Can you trust an AI model to do what you intended? This is a central question both for those deploying AI systems and for those seeking to evaluate their capabilities. In deployme…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Managed Deep Agents and LLM Gateway hit public beta, plus Deep Agents v0.7, Tuned Evaluators, Bring Your Own Cloud on AWS, and LangSmith Engine upgrades.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Managed Deep Agents and LLM Gateway hit public beta, plus Deep Agents v0.7, Tuned Evaluators, Bring Your Own Cloud on AWS, and LangSmith Engine upgrades.
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Amazon Bedrock AgentCore Evaluations decouples agent evaluation from the framework you build on. As long as your agent emits OpenTelemetry telemetry, the service can score it, whether you use LangGraph, LlamaIndex, the OpenAI Agents SDK, Google ADK, the Claude Agent SDK, or Strands Agents. This post explains how the framework-agnostic contract works.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Amazon Bedrock AgentCore Evaluations decouples agent evaluation from the framework you build on. As long as your agent emits OpenTelemetry telemetry, the service can score it, whe…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Google has updated Gemini Audio with some new Gemini 3.5 models, introducing new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Live, 3.5 Live Experimental, and 3.5 Transcribe are designed to provide better precision for Google's voice-controlled AI features, without struggling with background noise or when your speech is interrupted. Gemini 3.5 Transcribe is a completely new addition to the Gemini family, and its introduction comes as we're still waiting for Google to release the Gemini 3.5 Pro model that it promised to roll out in June. Google says that 3.5 Transcribe "repres … Read the full story at The Verge.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Google has updated Gemini Audio with some new Gemini 3.5 models, introducing new transcription capabilities that automatically detect specialized jargon and more than 85 languages…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Open-weight model runner Ollama has reintroduced an integration with Claude Desktop that lets users connect Anthropic’s app to models served The post Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled appeared first on The New Stack.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Open-weight model runner Ollama has reintroduced an integration with Claude Desktop that lets users connect Anthropic’s app to models served The post Claude Desktop can now easily…
AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Dive into LangSmith product usage patterns that show how the AI ecosystem and the way people are building LLM apps is evolving.
AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
Dive into LangSmith product usage patterns that show how the AI ecosystem and the way people are building LLM apps is evolving.