AI News HubLIVE

研究动态

待翻译:How to opt out of AI training on ChatGPT, Google, Meta and more

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Back to list Many AI features built into the apps you already use are improved using data from the people who use them. On several of the services in this guide, your prompts, posts, streams, or documents can end up sha…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Back to list Many AI features built into the apps you already use are improved using data from the people who use them. On several of the services in this guide, your prompts, pos…
站内正文

待翻译:Enhancing Agent Retrieval with Structured Chart Extraction

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The MotivationMore and more enterprises are now asking agents to work with their...

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The MotivationMore and more enterprises are now asking agents to work with their...
站内正文

待翻译:Show HN: Turn ad-hoc subagents into durable, accountable AI teams

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Notifications You must be signed in to change notification settings Fork 3 Star 78 BranchesTags Open more actions menu Latest commit History 135 Commits 135 Commits Folders and files NameName Last commit message Last co…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Notifications You must be signed in to change notification settings Fork 3 Star 78 BranchesTags Open more actions menu Latest commit History 135 Commits 135 Commits Folders and fi…
站内正文

待翻译:What We Can Learn From Google Engineers’ Indispensible Prompts

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Hey, Google Engineers: What prompt do you personally refuse to work without, and why?

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Hey, Google Engineers: What prompt do you personally refuse to work without, and why?
站内正文

待翻译:AI can detect heart disease in women using mammograms, study suggests

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Experts say findings mean breast screenings for cancer could also flag cardiovascular problems in women Doctors have discovered a way to use routine mammograms that screen for breast cancer to spot heart disease, the world’s leading and frequently underdiagnosed and cause of death in women. Researchers analysed the scans using artificial intelligence and were able to successfully identify women with coronary heart disease, high blood pressure or who had suffered a stroke. Continue reading...

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Experts say findings mean breast screenings for cancer could also flag cardiovascular problems in women Doctors have discovered a way to use routine mammograms that screen for bre…
站内正文

待翻译:OpenAI’s executive exodus has one big winner

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Today on Decoder, I’m talking to Verge senior AI reporter Hayden Field about some pure Decoder bait: the seemingly-endless org chart changes at OpenAI, and how all of them seem to consolidate power under cofounder Greg Brockman, the company’s president. While Sam Altman is the CEO and still OpenAI’s most public face, Brockman has amassed enormous power and influence within the top ranks of the company as other senior leaders have left in rapid succession these past few months. Verge subscribers, don’t forget you get exclusive access to ad-free Decoder wherever you get your podcasts. Head here. Not a subscriber? You can sign up here. Brockman now oversees the company’s entire consumer and enterprise product teams, including ChatGPT, Codex, and its major infrastructure build out. As Hayden recently reported, he is effectively now the day-to-day operational leader of OpenAI. This will have major effects on OpenAI as a company and its product strategy, at a time when it continues to cede ground to Anthropic in the enterprise and is preparing for a historic IPO. This is all against the backdrop of needing to turn a profit in the next few years and the company’s huge ambitions to replace both Google Search and the iPhone in the consumer market. So I wanted Hayden to break down for me what’s going on at OpenAI, and the increasingly important role Greg Brockman will play in its future. Okay: Verge senior AI reporter Hayden Field on Greg Brockman’s consolidation of power at OpenAI. Here we go. This interview has been lightly edited for length and clarity. Hayden Field, you’re The Verge‘s senior AI reporter. Welcome back to Decoder. Thanks. It’s great to be here. Always, always chaos when you’re here, Hayden. Absolutely. There’s never a calm week. If there’s a calm week coming up, I know something even crazier is coming the next week. We should just rename the show The Real Housewives of AI. [Laughs] Honestly, that would be fitting. Straight up, that’s what we should do. A lot of personalities, a lot of feelings, a lot of relationships that people have really come to value over a long period of time. And a lot of lore. A lot of lore that goes between all these people for years and years and years. It’s crazy. And now they’re even putting out profiles on some of their spouses. The circles and the people themselves are really interesting. Today the drama is about OpenAI and specifically Greg Brockman, who seems to be consolidating even more power at OpenAI. You just wrote a long story about this. It seems very clear that Greg is emerging as the central decision maker at OpenAI. Describe what’s going on. Greg has obviously been pretty influential at OpenAI for a really long time, but what’s different now is that he has so much control over the day-to-day operations in a way he didn’t before. He was always a cofounder. He was always heavily involved. It was him, Ilya Sutskever, and Sam Altman for a long time, all in these email threads that came out during the Musk v. Altman trial. We saw them all talking about the future of OpenAI, what it should look like, and all their strange dynamics with Elon Musk. But Brockman was in a big-picture role before, and now he’s increasingly taking on so much power in the day-to-day. Altman has been focusing more on the big picture, the IPO stuff, and the direction of the company at large. Meanwhile, Brockman is amassing all this power day-to-day as other executives leave. He’s been in control now of Codex, of its enterprise business, the consumer side. There are four arms of the company right now and he’s in control of all of them. It’s been really interesting to see how that’s happened over the last few months. Greg is a character in the AI story. Certainly inside of OpenAI, he’s been a central figure in a lot of the company’s drama over time. So if you’ve been paying attention to the AI industry, you’re familiar with him. But if someone haven’t, explain who Greg is and why he’s such a notable figure both in AI and at OpenAI. He has been on my radar for seven years now. He’s been on other people’s for even longer. He dropped out of MIT to join Stripe way back in the day. He was Stripe’s CTO during its really early explosive growth phase. He left in 2015 to cofound OpenAI. That’s how he came on the AI radar. Ever since then, he’s been a really influential figure. He, Ilya Sutskever, and Sam Altman were all heavily involved in the drama with Elon Musk really early on, raising money, trying to lure people over to OpenAI, hire top talent, and figure out who was going to control AGI in the event they ever made it. There’s a lot going on there. In the years following, Mira Murati joined and Brockman, Sutskever, Murati, and Altman became the four main players here. These were the most talked-about people at OpenAI. You saw them in the headlines the most. They had the most power at the company. And then when the board coup happened in 2023, that’s when Brockman flew onto the scene in a new way. He had been behind the scenes a little bit, under the radar. He was not really doing that many public interviews. I interviewed him in 2020, but it’s not like he was a talking head. You weren’t seeing that much of him unless you were really watching the company and the industry. But when Sam Altman got fired by the board, Brockman was so incensed that he quit immediately. He and Altman were going to start their own thing, or were potentially going to lead a department at Microsoft. That’s when he was making waves in a new way because he was hellbent on the drama. Brockman said, “Okay, if Sam goes, I go.” That’s when he became Sam’s right hand in a new way. Before that, he was just one of the many execs. They were close, but the board coup made them closer than ever. That’s also what made Sam know that he could trust him in a new way and thereby later greenlight him getting more power. This was a big bet in the politics of the organization. Mira was going to become the new CEO for five minutes. Then there was going to be another new CEO and it was unclear who would stay and who would go. Justifiably, Altman did not know who he could trust in that moment, except Greg Brockman was ride or die. He said, “I’m leaving. I’m going.” The five minutes that they were going to start a new division of Microsoft were some of the most hilarious five minutes in tech history. But Brockman had basically made it clear he was ride or die for Sam. Over time, that bet has paid off. That move has paid off. Even during that board coup moment, it was not clear that Sam would ever come back. I remember I was backpacking in Patagonia at the time and I didn’t have a laptop, but I wrote six articles on my phone. It was a crazy time. It seemed like Brockman and Altman were going to be fine, but just be on their own. Of course, all the hundreds of employees signing a letter that they would leave if Sam was not reinstated is what got the ball rolling on him coming back and then pulling an UNO reverse on all the board members who had voted to oust him, except one. That was a big moment where he didn’t know who he could trust. Some of the people who he had trusted the most were the ones who voted for him to be fired. That went a long way with him that Brockman quit as well. Ever since then, we’ve seen Brockman positioned in the public eye a little bit more. He’s been in the press more. This summer is when he really amassed a lot more power after a couple other executives left, and he’s now in control of quite a lot of pillars of the company. What’s really striking is, yes, we can talk about 2023 and the board coup and Sam going and coming back and who was going to stay or go then, but actually this year has been more dramatic in terms of executive departures from OpenAI. Just since April, the list is staggering. I’m just going to run down it: Bill Peebles who ran Sora; Kevin Weil, who was head of product at Instagram and then VP of OpenAI’s science arm; Fidji Simo, who was supposed to be head of product and was at one point the AGI chief, left on health leave and then just never came back; Kate Rouch, the Chief Marketing Officer; Srinivas Narayanan, the CTO of B2B; Brad Lightcap, who was the former CEO and the head of special projects; and then just recently Denise Dresser, the Chief Revenue Officer, who is an important character if your company’s about to IPO. That is a lot of people, and it seems like Brockman just took all of their roles as all of those people left. Is that what’s happening here? More or less. Some of them definitely affected his position less than others. For some of them, like CMO Kate Rouch, leaving didn’t really give him that much more power. But what was interesting is Simo, the ex-chief of AI at the company, is the one whose absence left a huge gap for Brockman to fill. He took over a ton of different aspects. As you just alluded to, a lot of the people who had just left the company had also recently changed roles. Through one way or another, a lot of these people leaving created a power vacuum that Brockman could then capitalize on. It’s not like he had this master plan of amassing power, but either way, he ended up with a lot of power. What is interesting is when Brad Lightcap transitioned roles, for example, one of these other people took over his COO responsibilities. Guess who it was? Dresser. Then she leaves. Now some of those responsibilities fall to Brockman. It didn’t all happen cut and dry, but in one way or another, people switching roles and then leaving the company, he ended up with a ton more power. I can’t even calculate the ratio really of how much power he ended up with. Now he’s in control — to give you some context — of the entire product strategy of the company, which is obviously incredibly important when you’re about to IPO and you’re getting a lot of pressure to turn a profit. He’s in charge of the company’s entire scaling arm, as well as core product and platform; all of consumer, which includes health, commerce, personal finance, ChatGPT, and other consumer-facing stuff; and core infrastructure, ads, data science, and growth. What isn’t he in charge of, really? You know what I mean? I want to talk about Fidji Simo for one more second here. She was the former CEO of Instacart, but before that she was a really high-ranking executive at Meta. She ran the Facebook app. She was in charge of a lot of monetization. When she came to OpenAI, it felt like her role was to turn the consumer version of ChatGPT into a product. There were lots and lots of people leaving Meta to go to OpenAI for a while, such that basically we had reported that every all hands at Meta was about, “What are you going to do about OpenAI? They’re taking all of our people.” You just saw this exodus of people from Meta going to OpenAI in various ways, led by Fidji Simo, in a moment where it felt like what they wanted to do was take on Google as a big consumer product supported by ads. It’s funny, Altman actually copped to that ambition on the David Senra’s show last week — here he is, giving Peter Thiel credit for saying OpenAI should compete directly with Google: Sam Altman: He’s like, ‘The power of this is the power of the Google text box, it’s a text box you can type anything into and it does the right thing. Clearly, the empty text box worked for Google so why don’t you just double down on that?’ That’s pretty straightforward and that’s what they did. Altman says they went super hard at it — being the empty text box, taking the Google business model, which is one of the best business models in internet history. That was a big bet, and they hired Fidji Simo to go run that bet. But she left, and that does not seem to be the emphasis anymore. The emphasis is on Codex, on enterprise, on competing with Anthropic. What’s left of that? Was that just a misfire and she had the wrong ideas? Is it still the plan? This seems like the question about OpenAI. Some of it has fallen by the wayside, which we’ve seen for a bunch of stuff that OpenAI has tried in the past year or two. They vowed to stop doing all their side quests, which is ironically something that Simo herself wrote in an internal memo. Several episodes ago, you were on and we just talked about the series of code reds at OpenAI. Right now they’re focused on the key revenue drivers, which are enterprise and coding. That’s really what they’re all-in on. They’re also, of course, really focused on building a “super app,“ which is supposed to incorporate all this stuff together, and making that super app better. Besides that, they’re trying to cut the fat and operate as a leaner business, which has been helped by some of these executive salaries being cut. They’re trying to look better on their balance sheet when they IPO. There’s a lot of pressure to do that, especially when SpaceX just IPOed at a crazy valuation and Anthropic is apparently gunning to exceed that. We’ll see if that actually happens. When they changed Simo’s title from CEO of applications to CEO of AGI deployment, that was the most Decoder thing of all. I looked at their org chart changing and I thought, “This isn’t going to last.” Over and over and over again, they keep reassigning people. And then as you say, they leave and we see one person take over those roles and consolidate power. We’ll come to the IPO in a second, but where is Sam Altman in all of this? It seems like he should be running OpenAI. He really likes focusing on the long-term stuff. The events of the last couple of years and all of the lawsuits have shown him that maybe he isn’t supposed to lead direct teams in a huge way. He has a few direct reports, but should he really be involved in the minutiae of the day-to-day operations? Maybe not. Altman has said before, and in blog posts that the company put out years ago, that this is the trend that they’ve been following. I remember a year or two years ago, the company would put out summer blog posts saying that different C-suite executives’ roles would be changing slightly, and Sam would be focusing a little bit on the longer term, on research. This is just a continuation of that trend, in my opinion. As we know — and this is a big Decoder thing as well — the more power you have at a company, the higher your title, the less you probably are involved in the actual day-to-day operations, especially at a tech company. This is another example of that. Sam has always been interested in the research and the long-term stuff. Brockman is an engineer. He was described really early on at the company as an “engineering workhorse that pushed to build scaled-up systems that would train the AI and make it work.” So he’s always been seen internally by employees as a can-do person. He makes things happen and he has the engineering training that is required to scale a system. So I could see investors being thrilled about this. I could also see Altman saying, “I know this guy’s in my corner.” Altman already cleaned out the board. He’s surrounded by people that are going to support him. He trusts Brockman. So he says, “I actually don’t really want to be involved in the operational stuff quite as much. I want to be focused on the direction of the company. So you handle that. You’re an extension of me. I can trust you. Go ahead.” That’s what I think is happening here. You’ve talked to Brockman before. I’ve met him. He just did a video with our friend Joanna Stern. You can see he’s pretty direct. He answers the questions. What kind of character do you think he’s going to be as a person who’s operationalizing OpenAI? Is this “get stuff done, make the numbers go up” situation? Or is this more expansive of an approach? It’s going to be interesting to see how he handles this much responsibility, especially this many disparate arms of the company. He’s a really can-do person. He has a lot of engineering training. He’s pretty respected within the company, but I do think there’s going to be a lot of natural tension here because consumer, enterprise, health, and personal finance are really, really intense departments that handle data extremely differently. And then you’ve got all the core infrastructure stuff and the behind-the-scenes stuff that makes everything work. We’re going to see a lot of natural tension arise, especially because this company has a limited amount of funds. Even if it’s a huge number, it’s still limited in some way, especially now compared to a few years ago. They have a limited amount of compute. That’s something that every executive at OpenAI has run into problems with — allocating the compute. The research side of things says they don’t get enough sometimes, or the product side says they need all of it. It’s going to be really interesting to see how he squares all this, especially as someone who’s probably trying to make everyone happy. He’s not going to be able to, and so he’s going to have to make some hard decisions. This brings us to the IPO, because the IPO is where those hard decisions have to pay off. There’s been a lot of talk about OpenAI playing catch-up to Anthropic, particularly in the enterprise and with coding. There was a Wall Street Journal story last week that said OpenAI’s revenue in Q2 lagged behind Anthropic’s. Obviously, Anthropic will tell you they’re profitable even though we haven’t actually seen their numbers and we don’t know how they calculate it. You’ve mentioned the IPO several times now. SpaceX just went public. Big IPOs are all the rage. What role is the IPO playing in all these changes? Is it focusing the company down? Is it making that equity pay off? Is it needing to raise more capital? What’s the shape of it? It’s absolutely behind a ton of these changes. It is normal to see a lot of changes in the C-suite ahead of an IPO. But some of the sources I spoke to within the industry, who study the way these IPOs typically happen in tech, said that the interesting and unusual thing here is how many C-suite execs left in such a short timeframe, because they know that looks bad for the company. It’s one thing if you need to cut salaries, make the balance sheet look a little bit different, make the company leaner. That’s normal. It is an easy way to do that, to not replace someone. We can only speculate how much each of these people’s salaries was. This is an easy way, cut and dry, to help out the balance sheet. But doing all of this in this amount of time reminds me of what we were talking about last time on Decoder about Google and DeepMind and how maybe Demis Hassabis didn’t leave when Jeff Dean did because they didn’t want to make the company look bad. Now at OpenAI, people are leaving left and right — sometimes within 24 or 48 hours. Brad Lightcap, for instance, just got a new job as special projects head. Then a couple months later he’s out, even after being at the company for so many years. Part of it’s probably just cashing out. You can make a lot of money if you’re an exec and the company’s about to IPO and you have a lot of stock options. But it is unusual, my sources said, that it’s happening in this short amount of time. The IPO is interesting for a number of reasons, but to me, the most clarifying is that OpenAI is up against Anthropic. There are numbers coming out, reported numbers, about Anthropic’s finances that make that company look pretty good. The company hasn’t really challenged them in public. So we have to assume that they’re happy that there’s reporting out there that the numbers look good for them. But all of that is enterprise. They’re selling Claude to big enterprises, to the government. It’s effective. Maybe their token prices have high margins. It’s unclear what the internals of Anthropic’s business look like, but the numbers that we can get look pretty good and they support a big IPO. There’s a lot of excitement around Anthropic for that reason. That’s Anthropic as an enterprise software provider. That’s their business. It’s a thing it’s focused on. It’s all they do. OpenAI has taken a lot of shots. You’re talking about cost-cutting ahead of an IPO. Are they shutting some things down? Are they closing the aperture on all the things they’re trying to do to compete with Anthropic ahead of this IPO? It’s interesting because I’ve seen a slight shift from OpenAI. Earlier this year, the company was saying, “Enterprise and coding. That’s what we’ve got to focus on. We have to make money. We have to compete with Anthropic.” Now, however, ahead of the IPO, OpenAI is getting a little bit of different advice. Yes, the company needs to focus on those things because obviously that’s where the money is. But now it’s also being told, “You also have to differentiate yourself from Anthropic. You can’t just be a copy that’s lagging behind. You have to be doing something different.” OpenAI is going to really lean into consumer and hardware as well. One attorney told me that for OpenAI to be successful in the public eye and in investors’ eyes, the company needs “hardware and consumer products to sell to consumers, and luckily, that’s something that Brockman has experience with.” It’s hard to do everything, but they’ve had to narrow a ton and then slightly widen again and say, “Consumer is what we’re really known for.” They’re saying, “We’re the Kleenex of AI right now because of how we’re known in the consumer world.” People say, “I’m going to ask ChatGPT,” or, “I ChatGPT’d it.” And they’re talking about just using AI in general. Maybe they weren’t actually using ChatGPT, but that’s the way they’re referring to things a lot of the time. So OpenAI knows they need to capitalize on that. That’s one way they can be different from Anthropic. Of course, Anthropic also offers that, but they’re much more known for enterprise and that is why they’re making so much money. But OpenAI is leaning into hardware. The company has Jony Ive. It has to not cut some of this stuff entirely because then it’s going to look like Mark Zuckerberg with Meta. It can’t go all-in on something and then just completely cut it, or they’re probably worried about being a laughingstock. You have to really cut the true fat, which they did with a lot of their side projects. But as far as consumer, hardware, enterprise, and coding, those are their main bets right now. We’re going to see them really double down on that heading into the IPO. Come on, Hayden. Aren’t we actually in the metaverse right now? Technically, you and I are in the Metaverse together right this second. It worked. Absolutely. It happened. Should have made an avatar. Technically we’re here in our bodies on the internet. Yep, that’s true. What argument can you have except that the metaverse definitely worked and Mark Zuckerberg was super right about all of it? The consumer business is really hard. The scale of the play you need to make the consumer business work is on the order of fully overtaking Google, which seems very challenging. In hardware, you have to replace the iPhone. If you do anything other than replace the iPhone, people still have their iPhones. And then you’re going to lose to Instagram every single time. Have they said anything about how they plan to do either one of these things? Because replacing Google with a Google-level monetization engine seems very hard. Replacing the iPhone — even if you have Jony Ive — such that people don’t still have their iPhones, seems very hard. Neither one of these things has been in any kind of focus for me at least. Like we talked about last time, I really am skeptical. Hardware is hard. I’ve seen way too many companies crash and burn when they try to make an AI hardware device. It’s going to look beautiful because Jony Ive is in charge. But as for how useful it’s really going to be, especially in the era of AI populism when there’s a huge backlash against using AI, I don’t know how excited people are going to be on a broad scale to have something visible pinned to them or in their ears or on their table where you can tell they are using AI. It’s going to be really exciting for a subset of people, but for the broader public, Meta’s glasses are still uncool, still being called pervert glasses. With AI hardware, you’ve got a steep hill to climb. Let’s talk about what happens next with OpenAI, because those are the challenges. What you have now is a new leader who is, as you said, running most of the company. Greg Brockman was just on CNBC last week. He basically defended the turnover. He said this: Greg Brockman: “I’d say fundamentally, we’re a very resilient organization… If you look over the years, there have been different eras where we have different sets of leaders in place. I’m constant, Sam is a constant. I think that we are stronger because of that resilience and diversity.” So this might be what Greg Brockman has to say: “All these people are gone, but I’m still here. Sam’s still here, and the company’s still the same.” Do you think it’s just the thing that he has to say? Or do you think there’s something more real about OpenAI as a company, where you have these two leaders who are clearly ride or die for each other and you can swap in and out all these other executives and the company will still have its own unique vision? He would like us to think the latter, but it’s more of the former. It’s never good when you have a bunch of people leaving at once. It’s never good when you’re constantly restructuring or reorganizing the company. It’s bad all the way down. Employees are probably feeling weird. They have a new person in charge. They’re being shuffled between different teams. Their teams are headed by one person and then their boss’s boss is someone else, which probably means their bosses are weird. It’s not great for productivity, especially heading into such an important time for OpenAI. Even if it’s a move that makes sense and is good long term, it’s still going to be weird for people in the short term. In some ways the latter is true in terms of Brockman and Altman having been at the helm from the start. Someone who’s been there for eight months who leaves is not going to have as much of an impact. But we’re seeing some people that have been there for a really long time leave as well, like Brad Lightcap. He was someone that I had tracked for years and years. All of a sudden he’s out, after switching roles, and now he says he’s going to start something new. He says he’s implying he’s still going to work with OpenAI in some way, shape or form. He and Altman had a really friendly exchange on X after he left, but a lot of it’s for show. They’ve got to make sure that people aren’t skittish about the company right now. It’s something he just has to say. But it’s also true that if there were two people that would be the most influential if they left, and luckily those two are still there, Brockman and Altman. You’ve mentioned the public antipathy towards AI several times now. There is just a lot of anti-AI sentiment out in the world. And it’s pretty politically coded, although it scrambled some political lines. Hating AI is pretty bipartisan. Hating data centers is pretty bipartisan. We’ve got a lot of reporting about that on the site. Brockman is into politics. He donated $25 million to MAGA Inc. He’s one of the top donors to Trump overall. Do you think him being so openly political will help or hurt him as he becomes more of a character, more of a visible leader of OpenAI? It’s going to help the company from the outside, because they’re trying to push a lot of stuff through during the Trump administration and they really want a voluntary regulation framework. They want to be in Trump’s ear in a good way. It’s going to hurt Brockman internally, because a lot of employees at tech companies that I’ve interviewed in the past year, no matter what company they work at, are incredibly angry if their CEO or their C-suite isn’t doing enough to speak out against some of the Trump administration’s decisions. Now, Brockman isn’t just not speaking out against them, he’s also supporting the admin with tens of millions of dollars. I could see this really hurting him from the inside. Maybe some people don’t take him seriously. Maybe we will see some departures. It’s going to help the company from the outside, just because when you’re giving a lot of money to Trump, he seems to let you curry favor with him. We’ll see. To be clear, we’ve seen Sam Altman stand next to Trump and announce data center projects. So whatever reputation OpenAI was going to have because of its leaders, it might already have, but the specific political giving seems new. And as we head towards the midterms, it seems like a new challenge for the leaders of the company to be openly associated with. Do you think it’ll change any of the valence around data centers and AI, how the public feels about them? I don’t know that it’ll change that because we’ve been careening towards this for a while. People are really mad about data centers and it’s a bipartisan feeling, like you mentioned, and same with AI. Sam Altman and other AI CEOs have said in the past month that they feel like AI has a big PR problem and that they haven’t done a good job of showing people the good parts of AI and how they need to do better to show people why they’re even building it in the first place. Maybe they do, but people have been told a lot of the good things about it and they’re still not thinking that they’re as good as the bad is bad. The AI industry is in for a rude awakening there. But as for what you mentioned about Altman, yes, he and every other AI CEO have been at Trump’s dining table. They’ve been caught on hot mics praising him. It’s not something hugely new that Brockman’s giving this money, but what is new is that he’s giving it so much of it in a personal capacity, so much so that Altman even had to come out and say at some point, “Brockman only did this in a personal capacity. We’re not saying that we’re super behind him.” He also couldn’t really say they weren’t behind him, but OpenAI had to separate themselves a little. That’s what’s interesting to me about this. It’s like the personal capacity of his giving has made so many headlines and OpenAI tried to distance itself a little bit, but not too much. Now they’re not going to be able to do that because he’s essentially running day-to-day operations. I’m very curious to see how that plays out. I don’t think we quite know yet, but as he becomes a more visible leader of the company, being that open about his politics, even in the context of the other big tech CEOs, is different. We don’t quite see that from all of them. We see it from Elon, but everyone else plays it pretty safe. This is new, especially for a leader at OpenAI. All right, let’s end here. Let’s say it’s two years from now, OpenAI is public. Do we think Greg Brockman is the CEO and Sam is just the “chairman of raising money,” or whatever it is that he’s doing? I could see that happening. It would take a couple of big things to make that happen. Altman really likes being in charge, so I don’t think he’s going to go super quietly. But if he ever needs to move on to a more big-picture chairman role, if the IPO doesn’t go so well or he gets dragged in the public eye or something and still needs to be part of the company but not the CEO, I could definitely see that happening. Especially because Brockman has a lot of ambitions. We saw his journal entries come out during discovery, during Musk v. Altman, and he was writing some pretty ambitious things. “What will get me to a billion dollars?” He’s saying, “This is my one chance to be in charge.” There are a lot of entries that he’s writing about how he’s put his blood, sweat, and tears into the company and how he needs to be in charge in some way, shape, or form. He will probably be gunning for a role like that and he’d be happy to take it on. We’ll see what happens in the next few years. But I definitely don’t think that’s out of the question, especially because he’s been there since the beginning. He has a lot of product experience and he’s clearly ambitious. The other thing that is interesting is, way back in the day, he wasn’t always so aligned with Altman all the time. In fact, Brockman and Sutskever, OpenAI’s chief scientist, were pretty aligned against Sam every once in a while. Not against him, but they were questioning him and pushing back. They were saying things like, “It seems like you really care about this CEO role. It seems like you really care about being in charge and having a lot of power in this way and that way. How does this relate to your political ambitions? How does this relate to who’s going to control AGI? We really don’t want it to be a dictatorship.” Altman would probably be hard-pressed to give up the CEO title when he fought so hard for it, even 10 or 11 years ago. But if he ever does, it seems like Brockman would be happy to take it on. I will end by stating my prediction that I’ve been making almost the whole year now. I don’t think we will end 2026 with OpenAI as the same kind of company as when the year started. Actually I would say, given all of this turnover, that prediction has already come true. Structurally, it is a very different company with very different goals than the year started. But I’ll just put that prediction to you. You cover the company way more closely than I do. Do you think there’s any way that OpenAI looks the same at the end of the year as it did at the top of the year? Absolutely not. When a company goes public, so many things change. Especially when the end of the year is only five months away. It’s a scary number of months away. It’s tomorrow. Let’s say maybe not by December 31st, but six months from now, 100 percent it’s going to look different and probably sooner than that. They’re going to have to make a lot of changes. They’re going to be listening to their investors in a new way and they’re going to be beholden to them in a new way. There are only certain parts of this industry that make money and OpenAI is not in the lead in those parts. They’re going to have to go all-in on that type of stuff. The type of research that they want to do to stay at the frontier, they can make a case for that. Money-wise, they can say, “Look, in order to stay in the lead, we have to do this long-term stuff a little bit.” But they’re not going to be able to go all-in on it all the time because it costs a lot of money — and they have a limited amount. We’re going to see a lot of the anxieties that we saw Brockman and Altman voice a year ago, two years ago. I’ve been in the room with them where they’re talking about their fear that they’re running out of compute and how they’re going to scale. Their whole job for the next six months, the next year, et cetera, is scaling. They’ve had so many concerns about this. They’ve had so many fears about how they’re going to do this with their limited compute, especially now that they’re going public and they’re beholden to investors in a new way. There’s no way the company is going to look the same. Hayden, this has been great. Thank you so much for being on Decoder again. We’ll have you back when another season of The Real Housewives of AI kicks off, which is probably going to be tomorrow at the rate we’re going. [Laughs] Absolutely. Thanks. Questions or comments? Hit us up at [email protected]. We really do read every email!

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Today on Decoder, I’m talking to Verge senior AI reporter Hayden Field about some pure Decoder bait: the seemingly-endless org chart changes at OpenAI, and how all of them seem to…
站内正文

待翻译:I didn't plan to let an AI manage my to-do list

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:I didn't plan to let an AI manage my to-do list Eddie (my AI agent) now manages my to-do list.1 I hadn’t planned that. It kind of happened on its own. As I wrote in one of my previous posts, I like interacting with him…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • I didn't plan to let an AI manage my to-do list Eddie (my AI agent) now manages my to-do list.1 I hadn’t planned that. It kind of happened on its own. As I wrote in one of my prev…
站内正文

待翻译:Piloting the first double-blind AI evaluations

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:August 27, 2026 Responsibility & Safety Piloting the world's first double-blind AI evaluations William Isaac, Sol Messing and Kristian Lum Share Building trust in proprietary model benchmarks using cryptographically sec…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • August 27, 2026 Responsibility & Safety Piloting the world's first double-blind AI evaluations William Isaac, Sol Messing and Kristian Lum Share Building trust in proprietary mode…
站内正文

待翻译:Fragmented AI Is Creating a "Faster but Not Better" Workplace

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Press Release New Workday Research: Fragmented AI Is Creating a "Faster but Not Better" Reality for Employees in Hong Kong and Taiwan Download PDF Around a quarter Hong Kong and Taiwan workers spend a significant amount…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Press Release New Workday Research: Fragmented AI Is Creating a "Faster but Not Better" Reality for Employees in Hong Kong and Taiwan Download PDF Around a quarter Hong Kong and T…
站内正文

待翻译:When agents act on their own, governance has to live in the data layer

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Presented by EDB As enterprises give AI agents more autonomy — the ability to plan, decide, and act across systems without a human approving each step — a hard question moves to the center of every architecture review: When an agent tries to complete an action that it was never authorized to do, what actually stops it? These are your agents, running on your models, touching your data in your infrastructure — and the responsibility for what they do sits with you. That responsibility can’t be met in hindsight or with a set of abstract policies that live on paper but not in practice. Agents need rules in the context of the moment, because they don’t exercise overriding judgment of their own actions. Consider a simple rule: Never open the car door. Followed literally, an agent could never get in or out of the car at all. But if you change the context (the car has just crashed, there’s a fire, someone is hurt and needs to get out), then the rule you actually want is the opposite. Context in the moment is everything. We are asking agents to do intelligent things; that requires intelligent rules. The instinct is to add guardrails around the agent: instructions, policies, and monitoring layered above the model. Those mechanisms matter, but they share a structural limit: The car-door rule is plausible right up until the moment you actually have to decide whether to open the door. Controls at the agent layer are only as reliable as the agent’s output is predictable, and autonomy is precisely the property that makes that output hard to predict. Governance that depends on reviewing an action before it happens cannot keep pace with a system that acts in milliseconds, across many systems at once. Governance has to become executable, and enforced where agents actually do their work: at the operational data layer, in the context, and exactly at the moment it is happening. The data layer is the enforcement point Agents create value by touching data. They query it, retrieve it, transform it, and increasingly act on it. A policy that says an agent should not reach a certain class of data is meaningful only if the system can deny that access at the moment the agent requests it. Additionally, a principle that says AI must be auditable is meaningful only if the organization can reconstruct what the agent did, what data it touched, which user it acted for, and what resulted. When governance lives at the data layer, it holds regardless of how the agent was built or how it behaves, because the control is a property of the database itself, not a promise made by the agent. Agent behavior may be probabilistic. Governance cannot be The enterprise should not rely on a model choosing to follow policy. The policy has to be enforced by the system. That is the difference between hoping an actor stays in bounds and constructing bounds it cannot cross to begin with. The controls that make this real are ones many enterprises already run at the data layer: role- and attribute-based access, row- and column-level security, classification and masking, policy as code, and complete audit trails. What agents change is not the mechanism, but who the mechanism has to recognize. Identity management has to treat the agent as a principal in its own right, with its own identity and a purpose declared when the session opens. Once purpose is bound to identity, the policy engine can evaluate it the same way it evaluates role or department today, and the record of what happened can capture not just who acted and what they touched, but what they declared they were there to do. In practice, this resolves into nine controls, grouped under three imperatives: Enforce it Role- and attribute-based access control enforced at query time, for agents as well as users Dynamic column masking driven by the same policy path Agent identity as a first-class principal, with declared purpose bound at session start and the acting user preserved See it and prove it Classification and tagging that drives policy Session-level audit logging that records which agent acted, for which user, and under what declared purpose Lineage across pipelines, so a result can be traced back to the request that produced it Unify and harden Centralized, portable policy management Encryption at rest and in transit Consistent enforcement across on-prem, cloud, and sovereign or air-gapped environments “Declared purpose is what makes the difference. It becomes an attribute the access layer already understands, evaluated in the same policy path as role and row-level security. The enforcement mechanism does not change. What changes is that the agent's purpose is part of what it evaluates, and part of what the record proves afterward,” says Priyanka Jain, VP, product management, data & AI governance, EDB. Wherever you are in your AI adoption journey, enforcement at the data layer is what lets you move faster rather than slower. The controls are already in the database. The difference is that agents now have to pass through them. A digital leash, not a locked door The goal is not to stop agents from doing useful work. It is to define how far an agent can go, what it can touch, what it can change, what requires escalation, and how the organization can reconstruct events if something goes wrong. Governed this way, agents are identified, scoped, monitored, and auditable. The enterprise can adopt them faster, because security, risk, and leadership teams trust the operating model underneath. Open, sovereign, and enforceable at the source Built on open source Postgres, this open foundation keeps enterprises in control of where their data lives, who can reach it, and under what policy, without ceding governance to a layer they don’t own or can’t inspect. For regulated industries, that combination of data sovereignty and source-level enforcement isn’t a nice-to-have; it’s the precondition for putting agents into production at all. Agentic systems will keep getting more capable and more autonomous. That is a reason to be deliberate about where control lives, not a reason to slow down. The enterprises that enforce governance at the data layer can move aggressively on AI, because the thing protecting their data is more than just wishful thinking. EDB Postgres AI is an open, enterprise-grade sovereign data and AI platform that unifies transactional, analytical, and AI workloads — with governance enforced where the data lives. For the full framework, see EDB’s white paper Governing Agentic AI at Enterprise Speed. Max Romanenko is Chief Technology Officer at EDB. Sponsored articles are content produced by a company that is either paying for the post or has a business relationship with VentureBeat, and they’re always clearly marked. For more information, contact [email protected].

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Presented by EDB As enterprises give AI agents more autonomy — the ability to plan, decide, and act across systems without a human approving each step — a hard question moves to t…
站内正文

待翻译:LetItLoop: Resume crashed AI agent loops with 0% token waste (<1ms resume)

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Notifications You must be signed in to change notification settings Fork 3 Star 4 BranchesTags Open more actions menu Latest commit History 205 Commits 205 Commits Folders and files NameName Last commit message Last com…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Notifications You must be signed in to change notification settings Fork 3 Star 4 BranchesTags Open more actions menu Latest commit History 205 Commits 205 Commits Folders and fil…
站内正文

待翻译:The Identity Crisis No One Planned For: Governing Non-Human Agents at Enterprise Scale

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:For a decade, identity and access management meant one thing: governing the humans who log in. Employee joins, gets provisioned, gets a manager, gets a departure date, gets offboarded. That loop is well understood. What changed is that the fastest-growing population inside enterprise environments is no longer human, and the governance playbook written for people […]

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • For a decade, identity and access management meant one thing: governing the humans who log in. Employee joins, gets provisioned, gets a manager, gets a departure date, gets offboa…
站内正文

待翻译:Conveo.ai (YC S24) Is Hiring – Senior Product Engineer NYC

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Senior Product Engineer ($220k-$300k + Equity) - NYC at Conveo | Y Combinator Conveo Confident decisions in days with AI-led interviews. Senior Product Engineer ($220k-$300k + Equity) - NYC $220K - $300K•New York Job ty…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Senior Product Engineer ($220k-$300k + Equity) - NYC at Conveo | Y Combinator Conveo Confident decisions in days with AI-led interviews. Senior Product Engineer ($220k-$300k + Equ…
站内正文

待翻译:The people, the money and the ownership behind China's leading AI companies

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:China & AI | WireScreen Briefings CareersProduct WWIRESCREEN · SPECIAL REPORT · CHINA AND ARTIFICIAL INTELLIGENCE WS-2026-034 · AUGUST 2026 Built and Owned The people, the money and the ownership behind China’s leading…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • China & AI | WireScreen Briefings CareersProduct WWIRESCREEN · SPECIAL REPORT · CHINA AND ARTIFICIAL INTELLIGENCE WS-2026-034 · AUGUST 2026 Built and Owned The people, the money a…
站内正文

待翻译:How decoding beluga whales’ chitchat may save them – and teach us more about ourselves

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:From humpbacks learning the latest songs to sperm whales’ ‘clan codas’, drones and AI learning are helping reveal another dimension to the concept of culture In 2015, the whale researcher Valeria Vergara pitched a small tent at the icy water’s edge on Somerset Island in the remote Canadian high Arctic, shrouded in a freezing fog so thick she struggled to see her own hands. She trailed a cord out to the coast and plopped a hydrophone into the water to record the calls of belugas. As many as 1,000 of the whales, accompanied by their newborn calves, frequented the surrounding bay. Belugas’ ceaseless chatter, made up of dozens of unique sounds, is crucial to keeping these creatures connected in the dim, turbid waters of the Arctic. Continue reading...

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • From humpbacks learning the latest songs to sperm whales’ ‘clan codas’, drones and AI learning are helping reveal another dimension to the concept of culture In 2015, the whale re…
站内正文

待翻译:The best iPad Air cases of 2026: Expert tested

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:The right case can take your iPad Air to the next level. These are our favorite iPad Air cases from brands like Apple, Burga, and Logitech.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • The right case can take your iPad Air to the next level. These are our favorite iPad Air cases from brands like Apple, Burga, and Logitech.
站内正文

待翻译:Two German airport workers die of malaria after 'mosquito arrives on plane'

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Two German airport workers die of malaria after 'mosquito arrives on plane' - BBC News Image source, Getty Images Image caption, Frankfurt is Germany's busiest airport ByAndré Rhoden-Paul and Joe Coughlan Published 26 A…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Two German airport workers die of malaria after 'mosquito arrives on plane' - BBC News Image source, Getty Images Image caption, Frankfurt is Germany's busiest airport ByAndré Rho…
站内正文

待翻译:GitHub – rajnandan1/ken: Thompson-mode systems discipline for AI agents

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 0 Star 1 BranchesTags Open more actions menu Latest commit History 18 Commits 18…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Uh oh! There was an error while loading. Please reload this page. Notifications You must be signed in to change notification settings Fork 0 Star 1 BranchesTags Open more actions…
站内正文

待翻译:What RTO Means for AI

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Disclosure: These views are my own and do not represent my current or any former employers. Executive summary The COVID-19 pandemic forced a large part of the North American knowledge workforce to work from home. The ch…

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Disclosure: These views are my own and do not represent my current or any former employers. Executive summary The COVID-19 pandemic forced a large part of the North American knowl…
站内正文

待翻译:Google Research Introduces GlucoFM: A 0.72M-Parameter Dual-Stream Foundation Model for Continuous Glucose Monitoring

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead of encoding it as one sequence. At 0.72M parameters it reached 58.8 task-averaged PR-AUC across 14 cohort–task evaluations, beating a 135M GluFormer and a 385M MOMENT. It remains a research prototype with no regulatory clearance. The post Google Research Introduces GlucoFM: A 0.72M-Parameter Dual-Stream Foundation Model for Continuous Glucose Monitoring appeared first on MarkTechPost.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • Google Research and UNSW Sydney released GlucoFM, a self-supervised foundation model that splits a CGM trace into a slow physiological stream and a transient event stream instead…
站内正文

待翻译:Generative Action-Chunk Sampling for Adaptive Stiffness Control in Physical Human-Robot Collaboration

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25284v1 Announce Type: new Abstract: Physical human-robot collaboration requires a robot to provide assistance when human intention is clear while remaining compliant when several future motions are plausible. We present an adaptive stiffness framework based on generative action-chunk sampling. Conditioned on an RGB image and external joint-torque estimates, the policy samples multiple future action chunks from an observation-conditioned prior. Variation among the sampled action chunks is used to continuously adapt joint stiffness and damping. Greater variation makes the robot more compliant to facilitate human guidance, whereas lower variation provides firmer assistance. In a real-world collaborative transport task with four possible directions, the proposed method achieved an average success rate of 0.95, compared with 0.83 for a fixed-stiffness ablation and 0.69 for a deterministic baseline. Near direction determination, variation among the sampled action chunks increased and the controller accordingly reduced stiffness. These results suggest that variation among actions sampled by a generative policy can serve as an online control signal for balancing assistance and compliance in physical human-robot interaction.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25284v1 Announce Type: new Abstract: Physical human-robot collaboration requires a robot to provide assistance when human intention is clear while remaining compliant w…
站内正文

待翻译:Development of a Voice-Controlled Tendon-Driven Bionic Hand

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25222v1 Announce Type: new Abstract: The impairment of the hands can seriously affect the abilities of every individual to perform the every-day activity, so the design of stable and controllable support devices is a significant field of study. This paper is about the design and implementation of an automated bionic hand which is dedicated to the coordinated finger movement through the simplified and efficient actuation mechanism. The method that the proposed system was designed on is the tendon-based method whereby the servo motors generate the movement of the fingers, with assistance of the angular control which is calibrated. An actuation is controlled by a microcontroller that will be programmed by use of an Arduino-based microcontroller to carry out programmed gestures that include open hand, fist, pinch and half flexion. It has an interface that is voice command enabled to make it easy to interact with a Bluetooth based sender receiver architecture which offers an option of executing trained commands which are immediately converted to finger actions. To explore the motions behavior, finger coordination and control response to the input, the behavior of the experiment system is tested. The actuation of the fingers was found to take a total of about 7-8 seconds to achieve full flexion of all fingers in a sequence. The system showed repetitive and constant motion throughout several actuation cycles without loss of any apparent tension or precision of control. There was a stable grasp of objects of different shapes and sizes, which implied consistent coordination between the fingers. These findings indicate that the proposed system offers predictable and steady control behavior and has a simple and efficient mechanical and control architecture.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25222v1 Announce Type: new Abstract: The impairment of the hands can seriously affect the abilities of every individual to perform the every-day activity, so the design…
站内正文

待翻译:Longitudinal Robot Learning from Demonstration with Care Providers in a Home Environment

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25196v1 Announce Type: new Abstract: Learning from demonstration (LfD) methods enable non-expert end users to teach robots novel skills without explicit programming. However most evaluations of the usability of LfD with non-experts has been conducted in controlled laboratory environments with a robotics experimenter present. In this work we identify non-expert end users' key barriers when teaching robots via demonstration without live robotics expert feedback in a home environment. In our human subjects experiment we support the non-expert end users through two forms of demonstrator guidance developed in prior work: pre-training and adaptive feedback. Towards the ecological validity of the evaluation, we conduct this experimentation over multiple visits, with a population of care providers. Finally, we propose to open source the resulting LfD dataset of care providers teaching a robot assistive tasks over multiple visits to a home environment.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25196v1 Announce Type: new Abstract: Learning from demonstration (LfD) methods enable non-expert end users to teach robots novel skills without explicit programming. Ho…
站内正文

待翻译:CRESSim-Neo: A Batched GPU Simulation Engine for Surgical Robotics and Robot Learning

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25192v1 Announce Type: new Abstract: We introduce CRESSim-Neo, a batched GPU simulation engine for surgical robotics and robot learning. CRESSim-Neo combines position-based simulation of rigid bodies, deformable tissues, fluids, and strands with batched rendering, surgery-specific sensing, and a GPU-resident data pipeline. The engine supports applications including tissue manipulation, fluid suction, suturing, cable-driven robots, and ultrasound image synthesis. Direct access to physics and rendering buffers enables GPU-resident robot learning and zero-copy PyTorch integration using DLPack. We demonstrate CRESSim-Neo across rigid-body, deformable-body, and fluid simulation tasks, including vision-based and surgical robot-learning scenarios. On an NVIDIA RTX 4090, the engine achieves up to 2.03 million environment steps per second for 8192 parallel CartPole environments, and scales to batched surgical scenarios involving tissue deformation, fluid interaction, and ultrasound sensing. Overall, CRESSim-Neo provides a unified and scalable platform for surgical simulation, synthetic data generation, and surgical robot learning.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25192v1 Announce Type: new Abstract: We introduce CRESSim-Neo, a batched GPU simulation engine for surgical robotics and robot learning. CRESSim-Neo combines position-b…
站内正文

待翻译:Control-Oriented Learning for Dynamic Tracking and Stability Analysis of Soft Pneumatic Actuators

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25171v1 Announce Type: new Abstract: Soft pneumatic actuators offer inherent compliance and safe interaction but remain difficult to model and control because of their highly nonlinear, distributed dynamics. We present a control-oriented data-driven modeling and control framework that decomposes actuator behavior into a nonlinear static equilibrium model and a linear residual dynamics model identified using Extended Dynamic Mode Decomposition with control (EDMDc). This representation enables feedforward compensation, task-space feedback control, and local closed-loop stability analysis through an augmented linear model. Experiments achieve approximately 1 mm root mean square error (RMSE) during low-speed (approximately 10 mm/s) trajectory tracking and below 10 mm RMSE at higher speeds (approximately 100 mm/s). The framework further achieves stable tracking of highly dynamic user-generated references with peak accelerations exceeding 25 m/s^2 while simultaneously performing real-time obstacle avoidance. Finally, the proposed stability analysis is experimentally validated by accurately predicting stable, marginal, and unstable operating regimes. These results demonstrate that structured, control-oriented learning provides an accurate and practical framework for soft actuator control.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25171v1 Announce Type: new Abstract: Soft pneumatic actuators offer inherent compliance and safe interaction but remain difficult to model and control because of their…
站内正文

待翻译:Sequential Object Placement Optimization with Convex Decomposition

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25162v1 Announce Type: new Abstract: Robotic object packing has been a core challenge for robotic deployment in logistics, industry, etc., due to the curse of dimensionality in combinatorial search and the difficulty of dealing with dynamic and contact constraints for irregularly shaped objects. Current heuristic and learning-based methods assume a limited spatial discretization resolution of space, and computation becomes extremely inefficient as discretization accuracy increases. In this work, we eliminate these assumptions by introducing SOPO-CD, a sequential optimization framework that frames object placement as a differentiable nonlinear optimization problem in a decomposed free space. We prove that placing a convex object inside a convex hull is essentially constraining the vertices of the object inside the convex hull. The constraints and their derivatives can be written in closed form and calculated within $200$ns. We implement a custom solver that achieves optimal placement within tightly constrained space in milliseconds; a $100 \times$ speedup compared to a classical grid search method. We generalize our framework to 2D Tangram, 2D Tetris, and 3D Bin Packing, and have demonstrated strong computational performance and packing utility. We also demonstrate solving a real-world Tangram puzzle online using an Allegro Hand and an Xarm.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25162v1 Announce Type: new Abstract: Robotic object packing has been a core challenge for robotic deployment in logistics, industry, etc., due to the curse of dimension…
站内正文

待翻译:SkyDrive: Learning to Drive in a New City from Aerial Traffic Monitoring

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25142v1 Announce Type: new Abstract: Autonomous driving has made remarkable progress through imitation learning with massive human demonstration data. However, a trained planner often degrades severely when applied to a new environment zero-shot, because of domain shifts in traffic regulations, road layout and driving behaviors. Therefore, adapting a trajectory planner to a new city typically requires resource-demanding local data collection with a vehicle sensor suite. In this work, we show that driving behavior can be learned from a scalable and efficient alternative. We introduce \emph{SkyDrive}, a framework that utilizes drone-based traffic monitoring to provide efficient supervision for autonomous driving agents in a new environment. While vehicle-based data collection logs the ego and its surroundings, an aerial platform naturally observes many road users simultaneously over an extended field of view. As a result, every vehicle can be a data source with grounded driving behavior, effectively scaling up the amount of supervision. Based on 137 hours of aerial traffic monitoring footage, we extract 650K driving samples and construct a benchmark for trajectory planners and motion predictors. Zero-shot experiments with multiple models reveal significant cross-city domain gaps, but many of them can be alleviated by limited supervision from the sky, e.g., 30 minutes of monitoring per location. Our findings show that aerial traffic monitoring is an efficient and scalable data source for adapting autonomous driving systems in new cities. Data and code will be made publicly available.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25142v1 Announce Type: new Abstract: Autonomous driving has made remarkable progress through imitation learning with massive human demonstration data. However, a traine…
站内正文

待翻译:Extending Ground-Constraint LiDAR-IMU Calibration to Tilted Surfaces in a Continuous-Time Framework

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25135v1 Announce Type: new Abstract: This paper presents a novel method that extends targetless LiDAR-IMU calibration for ground vehicles to non- flat environments. Calibration typically necessitates full exci- tation of the sensor rig, a requirement that is not fulfilled by ground vehicles in normal operation. To address the degenerate planar motion, state-of-the-art methods propose residuals that assume the colinearity of the gravity and physical surface normal vectors, restricting usage to cases where the ground is assumed flat. This paper proposes ground-plane residuals that do not require this assumption, and are applicable for planar motion on a tilted surface. Results are demonstrated on a dataset collected from a Husky ground vehicle, on the M2DGR dataset, as well as on an offroad vehicle dataset. Repeatability is shown to be improved both in tilted and flat-ground scenarios, with strong improvement demonstrated for the tilted case. The implementation and experiments are open-sourced at https://github.com/vkorotkine/licalib_tilted_ground.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25135v1 Announce Type: new Abstract: This paper presents a novel method that extends targetless LiDAR-IMU calibration for ground vehicles to non- flat environments. Cal…
站内正文

待翻译:ROS2 Connect: A new ROS2 over WAN Solution

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25102v1 Announce Type: new Abstract: The Robot Operating System 2 (ROS2) has become a widely adopted framework for the development of distributed robotic systems. However, its communication architecture, based on DDS and RTPS, relies on multicast discovery mechanisms that are typically unavailable in wide-area network (WAN) environments, making remote operation challenging. This work presents ROS2 Connect, a WebSocket-based communication framework that enables transparent and secure ROS2 interaction across routed networks without requiring modifications to network infrastructure or DDS configurations. The proposed client-server architecture supports bidirectional exchange of topics, services, actions, and system data while integrating authentication and access control mechanisms. Experimental evaluation over a real WAN connection demonstrates significantly lower latency, higher stability, and improved scalability compared to existing solutions, including DDS Router, rosbridge and Zenoh. Initial results show that ROS2 Connect provides a reliable foundation for teleoperation and distributed robotics applications over wide-area networks.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25102v1 Announce Type: new Abstract: The Robot Operating System 2 (ROS2) has become a widely adopted framework for the development of distributed robotic systems. Howev…
站内正文

待翻译:GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24959v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models encode visual observations as flat 2D patch tokens that carry no intrinsic geometric structure, and augmenting them with dense monocular depth injects per-pixel scalar values that encode neither surface orientation nor geometric confidence. This leaves the policy with limited structured spatial reasoning for action prediction. We propose GaussVLA, a Mamba-based VLA that incorporates two custom modules: Gaussian Spatial Tokenizer (GST) to lift frozen semantic and depth features into compact 3D Gaussian tokens, pools geometrically salient regions with learned queries, and \emph{Depth-Aware Chain-of-Thought (DA-CoT)} that performs structured, non-autoregressive geometric reasoning under language and flow-time conditioning. Across both simulation and real-world evaluations, GaussVLA demonstrates strong spatial-manipulation performance while remaining parameter-efficient. On LIBERO, it achieves 93.5% average success and 100.0% success on the Spatial suite with only 200M parameters, improving over SpatialVLA by 19.7% relative average success while remaining significantly more parameter-efficient.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24959v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models encode visual observations as flat 2D patch tokens that carry no intrinsic geometric structure,…
站内正文

待翻译:Lowering the Barrier to AI-Driven Inspection: A No-Code Workflow for Automated Structural Defect Detection

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25176v1 Announce Type: new Abstract: Structural health monitoring (SHM) is essential in modern engineering, providing data for condition-based maintenance, lifecycle assessment, and predictive decision-making. Traditionally, SHM relied on visual inspection to detect defects such as cracks and deformations. Early computer vision (CV) methods, including thresholding, edge detection, and handcrafted features, aimed to automate this process but were highly sensitive to noise, imaging variations, and multiscale defects, limiting their reliability. Recent advances in machine learning, particularly convolutional neural networks (CNNs) and You Only Look Once (YOLO), have improved defect detection accuracy and enabled real-time analysis. However, adoption in SHM remains limited due to technical barriers such as data labeling, model training, and deployment, which typically require programming expertise. To address this gap, we introduce YOLOEZ, an open-source, GUI-based tool for end-to-end YOLO model application. YOLOEZ integrates data labeling, training, and inference into a single interface, enabling high-performance model development without code while supporting reproducible workflows. Evaluation against existing software and classical image processing demonstrates that YOLOEZ not only outperforms traditional methods across most detection metrics, but also lowers adoption barriers present in other modern CV tools. By combining accuracy with accessibility, YOLOEZ facilitates wider use of AI-driven monitoring for predictive maintenance, digital twins, and intelligent structural systems.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25176v1 Announce Type: new Abstract: Structural health monitoring (SHM) is essential in modern engineering, providing data for condition-based maintenance, lifecycle as…
站内正文

待翻译:See More, Detect Less? Taming Information Leakage in Multi-View Anomaly Detection

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25168v1 Announce Type: new Abstract: In multi-view anomaly detection, more cross-view information can actually hurt. When multiple inspection views are naively fused in a reconstruction-based pipeline, normal cues from intact views propagate to the decoder, which faithfully reconstructs anomalous regions, collapsing the reconstruction gap the detector depends on. We call this failure mode \emph{cross-view information leakage} and show that effective multi-view fusion must explicitly restrict the information reaching the decoder. Building on this insight, we present GLAD(Global-Local Attention Driven framework), the first framework combining vision foundation model features with local and global cross-view fusion for multi-view anomaly detection. The Multi-view Merging Attention (MMA) module performs local cross-view fusion at linear complexity with learnable view importance weighting and token-wise gating, letting each view selectively incorporate fine-grained evidence from other views at $\mathcal{O}(N)$ cost. The Object-Guided Attention (OGA) module captures global context by aggregating class tokens from all views into a single object-level representation and broadcasting it back to patch tokens via temperature-scaled sigmoid gating, replacing the original patch representations rather than adding a residual to preserve the reconstruction gap. Experiments on Real-IAD and MANTA-Tiny show that GLAD outperforms state-of-the-art methods across sample-, image-, and pixel-level metrics, confirming that principled information restriction is key to multi-view anomaly reasoning.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25168v1 Announce Type: new Abstract: In multi-view anomaly detection, more cross-view information can actually hurt. When multiple inspection views are naively fused in…
站内正文

待翻译:What Do Audio-Visual Synchronization Metrics Actually Measure?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25157v1 Announce Type: new Abstract: Automatic AV-sync metrics are widely used to rank and train audio-visual generators, but they are rarely audited as measurement instruments. We jointly audit AV-Align, ImageBind AV-relevance, JavisScore, and Synchformer/DeSync under a common reliability protocol: controlled-distortion monotonicity, preprocessing sensitivity, rank uncertainty, cross-metric agreement, PEAVS-proxy agreement, and learned fusion. The result is an axis split, not a single winner: Synchformer/DeSync is the strongest temporal-offset tracker ($\tau=0.84$), ImageBind/JavisScore better match the PEAVS human-aligned proxy ($\tau=0.20$) and content-disruption families, and AV-Align is the weakest standalone metric. The metrics mutually disagree (Krippendorff $\alpha=0.066$), and neither linear nor simple $k$-NN fusion improves PEAVS agreement over the best individual metric. We recommend reporting AV-sync as a Reliability Card (metric-family breakdowns with confidence intervals) rather than a single bare synchronization score.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25157v1 Announce Type: new Abstract: Automatic AV-sync metrics are widely used to rank and train audio-visual generators, but they are rarely audited as measurement ins…
站内正文

待翻译:Can You Trust Frozen Hematology Foundation Models under Acquisition Shift?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25148v1 Announce Type: new Abstract: Frozen hematology foundation-model (FM) embeddings reach near-saturated in-domain white-blood-cell (WBC) accuracy, but clinical deployment demands reliability across scanners, sites, stains and preparation pipelines. We audit 15 frozen encoders (hematology, pathology, and general vision) across four public single-cell acquisition domains along two axes: accuracy robustness and calibration. In-domain linear-probe macro-F1 is saturated (0.98-0.997), yet cross-dataset macro-F1 drops 34-72% and rankings re-order: DinoBloom-L, the in-domain best, falls to 10th of 15 on the most-shifted target (MLL23) at the benchmark's shared 224-px input, behind RedDino and several general and pathology encoders. Rank transfer is probe-dependent: 1-NN retrieval is more stable on average than a source-fitted linear head (median $\rho$ 0.65 vs 0.45), but neither probe universally predicts target robustness. Calibration also collapses: source-trained probes are nearly calibrated in-domain (expected calibration error, ECE, 0.004) but confidently wrong off-domain (ECE 0.35), and source-fitted temperature scaling transfers poorly. We further audit pretraining exposure and identify MLL23 as DinoBloom's internal cohort; because DinoBloom's only held-out dataset is also our source domain, this benchmark cannot isolate exposure from scanner-associated shift. Label-free adaptation and marginal-entropy-based model selection appear safe under balanced evaluation but fail under realistic WBC class-prior shift. Class-Balanced Re-standardization (CBR), a training-free pseudo-label-balanced feature normalization, improves all evaluated target-prior scenario means and partially improves calibration, although encoder-level exceptions and residual miscalibration remain. Hematology FM benchmarks must therefore jointly audit accuracy, calibration, exposure, and class-prior robustness.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25148v1 Announce Type: new Abstract: Frozen hematology foundation-model (FM) embeddings reach near-saturated in-domain white-blood-cell (WBC) accuracy, but clinical dep…
站内正文

待翻译:RefLAM: A Reference-Grounded Line Annotation Pipeline for Historical Arabic Manuscripts

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25140v1 Announce Type: new Abstract: Existing approaches to building line-level Arabic handwritten-text-recognition (HTR) training data either rely on fully manual annotation, which does not scale, or on automatic OCR-to-reference alignment methods not yet extended to multi-script, two-zone (main-plus-margin) manuscript layouts with a provable correctness guarantee. We present RefLAM (Reference-grounded Line Annotation for Manuscripts), a pipeline converting manuscript page images and clean transcriptions into validated, line-level ground truth without sacrificing human oversight. RefLAM couples a deep-learning page-segmentation model with a multimodal large language model (MLLM) for structured OCR and a diacritic-agnostic fuzzy alignment engine that grounds each OCR line in a contiguous span of the reference text, with a character-level confidence score in $[0,100]$. A perfect score is provably equivalent to character-for-character identity of the normalised strings (the Confidence-100 rule), verified with no counterexample across the released corpus. A reviewer can thus trust a perfect score, confirming most lines at a glance rather than retyping them, so annotation becomes triaged, with attention concentrated on uncertain alignments. Across 7 fully page-validated books we measured a 75$\times$ throughput gain over manual annotation (3,000 vs. 40 lines/hr); applying the same guarantee to 7 further books, we retained 16,533 confidence-100 main-text lines within one week, excluding sub-100 lines rather than manually correcting them. Using RefLAM, we release AraMS-28k: 14 historical Arabic manuscript books, 3,043 pages, and 27,971 main-text and 629 margin-line annotations with bounding boxes, layout labels, and insertion anchors for 191 margin entries (30.4%). We also finetune Muharaf-pretrained baselines (including HATFormer) on AraMS-28k and report CER results confirming its practical utility for downstream HTR training.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25140v1 Announce Type: new Abstract: Existing approaches to building line-level Arabic handwritten-text-recognition (HTR) training data either rely on fully manual anno…
站内正文

待翻译:SHIFT-LLM: Distribution Shift Correction in Depth-Pruned LLMs

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25068v1 Announce Type: new Abstract: Depth pruning removes entire Transformer blocks to reduce the inference cost of large language models, but disrupts the hidden-state distributions expected by downstream layers, leading to significant accuracy loss. We introduce SHIFT-LLM, a training-free post-pruning correction framework that inserts a Linear Residual Adapter (LRA) at each pruning site. Each LRA preserves the identity pathway of the original residual block and adds a lightweight affine residual correction. This correction is calibrated via closed-form least-squares regression on a small held-out set, without gradient computation, to approximate the missing residual update produced by the pruned block. Together with the preserved identity pathway, the resulting LRA output approximates the hidden state produced by the original block, thereby mitigating the distributional mismatch introduced by layer removal while avoiding the expensive attention and feed-forward computations of the removed blocks. The resulting LRAs support low-rank factorization and exact merging across consecutive pruned layers for additional compression, and combine naturally with parameter-efficient fine-tuning for further recovery beyond fine-tuning the pruned model alone. Experiments on five model families, six layer-selection criteria, and seven zero-shot benchmarks show that SHIFT-LLM consistently recovers accuracy lost to depth pruning across most configurations, achieving gains up to +15.7 points on Llama-3.1-8B-Instruct while requiring only a few hundred calibration samples and no gradient computation.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25068v1 Announce Type: new Abstract: Depth pruning removes entire Transformer blocks to reduce the inference cost of large language models, but disrupts the hidden-stat…
站内正文

待翻译:Targeting the Attention Heads Behind Object Hallucination in LLaVA

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24966v1 Announce Type: new Abstract: Vision-language models such as LLaVA-1.5-7B often hallucinate objects absent from the image when generating captions. We ask whether an interpretability diagnosis of this failure can guide a targeted fix, and we measure what that fix actually changes. We rank attention heads by how much their image attention drops around hallucinated object words, then screen the shortlist by ablating candidate heads and measuring the change in hallucination-token log probability, yielding a 32-head set. We restrict two interventions to these heads: a head-sliced LoRA adapter and an inference-time grounding controller. On 400 held-out COCO images, the combined method lowers CHAIRs (the fraction of captions with a hallucinated object) from 0.370 to 0.230 and CHAIRi (the fraction of hallucinated object mentions) from 0.156 to 0.096 (p < 0.001, paired sign-flip tests). Two controls sharpen attribution. A random-head LoRA control, matched layer-for-layer and trained identically, performs no better than the matched baseline on a separate 200-image control split, supporting the role of head selection rather than LoRA capacity. Under fixed decoding budgets, the CHAIR reduction persists and grows with budget (23% at 64 tokens to 58% at 128), arguing against a pure max-token or truncation artifact, although the method remains shorter and more conservative. The resulting behavior reduces unsupported object mentions while also lowering object recall (0.78 to 0.70). We present a diagnosis-to-intervention pipeline for object hallucination, and, more importantly, a controlled account of what acting on the diagnostic signal actually does: it localizes intervention sites with real, non-random leverage, reported as a behavioral profile rather than a single score.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24966v1 Announce Type: new Abstract: Vision-language models such as LLaVA-1.5-7B often hallucinate objects absent from the image when generating captions. We ask whethe…
站内正文

待翻译:Synergising Local Geo-Environmental Characteristics with Spatial Context for Enhancing Landslide Susceptibility Mapping

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24956v1 Announce Type: new Abstract: Data-driven methods are widely used in landslide susceptibility mapping (LSM) because they can effectively model the complex relationships between landslides and geo-environmental conditions. Existing data-driven approaches generally follow two types of data representations. Pixel-based models focus solely on the geo-environmental characteristics of a specific landslide but neglect the influence of its surrounding environment. Patch-based models incorporate surrounding spatial context but may include pixels with weak or no spatial relevance to the target landslide location. To address this limitation, this study proposes a Local-Geo and Spatial Context Fusion (LGSCF) strategy, which synergises the geo-environmental characteristics of landslide points with their corresponding spatial context through a feature-wise modulation mechanism. We tested the LGSCF strategy by integrating it into several representative convolutional neural network (CNN) architectures, creating nine different LGSCF-based models. The study area covers approximately 2644 km2 across Jenai and Sinyi Townships in Nantou County, Taiwan, and the dataset comprises 5332 landslide samples and an equal number of non-landslide samples. The results show that LGSCF-based models consistently outperform their original versions, achieving F1-scores up to 87.09% and AUC values up to 0.9472. Furthermore, the susceptibility maps produced by LGSCF-based models show that known landslides are more accurately concentrated in "very high" susceptibility zones with fewer misclassifications. These findings demonstrate that our fusion strategy can significantly improve the accuracy of landslide susceptibility mapping.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24956v1 Announce Type: new Abstract: Data-driven methods are widely used in landslide susceptibility mapping (LSM) because they can effectively model the complex relati…
站内正文

待翻译:A Lightweight Multimodal Vision-Language Framework for Early-Stage Anatomical Green Fruit Classification in Commercial Orchards

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24935v1 Announce Type: new Abstract: Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is essential for robotic thinning, crop-load management, and other precision orchard operations. This study presents a lightweight multimodal vision-language framework that adapts TinyCLIP for fine-grained fruitlet anatomy classification in complex orchard environments. A dataset of 600 high-resolution RGB images collected from Scilate and Scifresh apple orchards was converted into 224 x 224 image patches and annotated for three anatomical classes. Domain-specific language prompts, such as ``a photo of a class,'' were used to guide multimodal alignment between orchard imagery and horticultural structures. A sliding-window inference strategy with a stride of 112 pixels aggregates patch-level predictions into spatial heatmaps, enabling interpretable whole-image localization of fruitlet components relevant to robotic thinning. Patch-level evaluation on an NVIDIA T4 GPU achieved F1-scores of 0.95 for calyx, 0.98 for fruitlet, and 0.85 for peduncle, with a macro-F1 score of 0.93. Deployment-oriented optimization using ONNX and TensorRT enabled efficient inference on NVIDIA Jetson hardware, preserved accuracy under INT8 quantization, and supported model sizes of approximately 127-137 MB with millisecond-level patch inference. These results demonstrate that lightweight vision-language models can provide interpretable and edge-deployable perception for automated fruitlet analysis and future robotic thinning systems. The source code and implementation details are publicly available at https://github.com/WilliamBu1/A-Lightweight-Vision-Language-Model-for-Early-Stage-Fruitlet-Classification-in-Apple-Orchards.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24935v1 Announce Type: new Abstract: Accurate identification of early-stage apple fruitlet anatomical structures, including the calyx, fruitlet body, and peduncle, is e…
站内正文

待翻译:Fusing Perceptual Vision Experts with Multimodal Large Language Models for Explainable Plant Disease Diagnosis: From Benchmark Imagery to Real-World Robotic Field Validation

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24934v1 Announce Type: new Abstract: Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hybrid Hierarchical Multi-Agent Framework (H$^{2}$MAF), combining decision-level fusion of EfficientNet-B3 and ConvNeXt-Tiny with semantic arbitration by open-weight multimodal large language models (MLLMs), Gemma 4 E4B and Qwen3.5 4B, using structured JSON evidence to generate explainable diagnoses, risk levels, treatment urgency, and financial exposure. (H$^{2}$MAF) is evaluated on 14,364 images (1,370 test images) across PlantDoc (2,922 images, 27 classes) and two non-public, continuously captured Cornell robot-acquired field datasets: Stage 2 (20 GB; 4,215 images) and Stage 4 (40 GB; 7,227 images), covering Early Blight, Late Blight, and Septoria Leaf Spot under uncontrolled field conditions. On PlantDoc, Gemma improves accuracy from 63.9% to 68.5%, achieving +7.6 points on the 41.7% CNN-conflict subset. Cornell accuracies reach 99.3% and 98.9%, with only 1.7-4.1% disagreement, demonstrating conflict-dependent MLLM utility. The critical-risk error of gemma is 0.14-0.5 points, whereas Qwen overflags by 3.5-14.4 points. These results establish MLLM arbitration as a promising, yet calibration-dependent, approach for explainable agricultural AI and robotic field decision support. Github Link: https://github.com/Applied-AI-Research-Lab/Explainable-AI-Plant-Disease-Detection

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24934v1 Announce Type: new Abstract: Accurate field plant disease diagnosis requires reliable fusion of uncertain and conflicting perceptual evidence. We present the Hy…
站内正文

待翻译:DataKernelBench: Can LLMs Optimize Database Queries on GPUs?

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Existing LLM kernel benchmarks focus on machine learning operators, leaving irregular, heterogeneous, data-movement-heavy database-style operators untested. We introduce DataKernelBench, which translates SQL into validated PyTorch TorchPlan programs and evaluates LLMs that optimize either the core tensor-bounded snippet or the full query in CUDA or Triton through execution-guided repair. Across ten proprietary and open-weight models on TPC-H SF10 with an H100 GPU, the strongest full-query CUDA configuration achieves $2.11\times$ speedup over torch.compile at full pass rate. We find that higher-performing implementations commonly use kernel fusion and execution-strategy changes, stronger models benefit most from full-query specialization, and workload context matters more than hardware context. To handle data larger than GPU memory, we extend TorchPlan with Dask-cuDF for on-demand partition loading on TPC-H SF100 with four H100 GPUs, achieving $2.54\times$ speedup

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25061v1 Announce Type: new Abstract: GPUs increasingly accelerate database systems, but query-specific peak performance still often relies on hand-written kernels. Exis…
站内正文

待翻译:Padamitra: Grounded Glossary Generation for Classical Sanskrit

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25038v1 Announce Type: new Abstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases and produce translation-grounded meanings from a sloka-translation pair, formalizing the traditional patha commentary practice as an evaluable NLP objective. We construct a benchmark of 31,316 sloka-translation-glossary triples from the Valmiki Ramayana and Srimad Bhagavatam, paired with two metrics: Jaccard for phrase recovery and Meaning Faithfulness for semantic consistency. Across zero-shot, few-shot, and instruction fine-tuned variants of Gemma-3n-E4B, Gemma-3-12B, Phi-4, and Qwen3.5-9B, instruction fine-tuning substantially outperforms prompting, while explicit segmentation yields gains. Error analysis identifies over-segmentation of sandhi and samasa compounds as the dominant failure mode, pointing to morphological modeling as the key bottleneck for faithful Sanskrit lexical decomposition.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25038v1 Announce Type: new Abstract: We introduce grounded glossary generation, a structured task requiring models to recover semantically meaningful Sanskrit phrases a…
站内正文

待翻译:Behind the [MASK]: Disentangling Representation and Faithfulness in DAPF-Based Dementia Detection

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25028v1 Announce Type: new Abstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia screening, but such models remain internally opaque. We study the interpretability of the Domain-Adapted models via Prompt-based Fine-tuning (DAPF) framework, which casts dementia detection as diagnosis-related masked-token prediction. We interpret DAPF and strong baselines using a variety of probing and analysis techniques, finding that DAPF achieved the best overall performance (accuracy=0.83 and macro-F1=0.83) with diagnosis most recoverable from its [MASK] representation. However, this representational advantage did not extend to token-level explanation faithfulness. DAPF attributions primarily reflected language task vocabulary, discourse markers, and transcription artifacts, with perturbation tests showing weak or negative effects. This suggests that its masked-token interface determines diagnosis information without producing faithful token-level explanations.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25028v1 Announce Type: new Abstract: Spoken-language analysis via prompt-based domain-adaptive models is a promising direction for low-resource, non-invasive dementia s…
站内正文

待翻译:A Primer on Computational Semantics for Artificial Intelligence Systems

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25022v1 Announce Type: new Abstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important to know how such models learn and represent the meaning of the language, and to be more informed about what language is. This document is an attempt to help the reader understand how linguistic meaning (i.e., semantics) is approached from different fields of scientific and philosophical examination. I also explain three primary semantic theories: formal semantics, grounded semantics, and distributional semantics then compare how transformer-based language models differ from how humans learn language.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25022v1 Announce Type: new Abstract: As people adopt transformer-based language models (e.g., ChatGPT and Gemini) for an increasing number of use-cases, it is important…
站内正文

待翻译:The Imperfective Paradox Is Not Necessarily in Large Language Models: A Benchmark Failure Before a Model Failure

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and reports that models frequently infer completed telic events from progressive descriptions, attributing this behavior to a Teleological Bias. It further argues that prompting interventions cause a Calibration Crisis. We reexamine the benchmark and conclusions and show that it is substantially affected by conceptual and evaluation mis-specifications. We identify three conceptual mis-specifications. In particular, Aspectual Reduction affects the benchmark construction, analysis, experiments, and conclusions. Under a strict NLI standard, 76% of Group A instances do not explicitly rule out culmination. In our native-speaker annotation, 38% of Group A examples and 29% of the Group C examples were judged to permit an alternative interpretation. To control these issues and lexical variation, we construct Lexically Matched Minimal Pairs. At the evaluation level, we formulate event-semantic NLI as a Multi-step Reasoning Problem and assess both intermediate semantic decisions and final predictions. Our results show that models often do not affirm culmination but nevertheless accept the corresponding simple-past hypothesis, a pattern we characterize as Sufficiency Bias. We further show that prompting interventions produce a Decision Shift among labels without reliably improving the underlying semantic understanding and reasoning. Intermediate and oracle-guided analyses identify two additional failure modes: errors in compositional aspectual classification and Surface-form Attraction toward surface-associated answers. Our experiments on Qwen-7B with suitable prompts, GPT-5.4, and Qwen-72B provide initial evidence for the context sensitivity of aspectual classification and suggest that these models can achieve performance comparable to that of human annotators.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.25005v1 Announce Type: new Abstract: The imperfective paradox provides a useful test of compositional semantic analysis. Recent work constructs an NLI benchmark and rep…
站内正文

待翻译:Does Fine-Tuning Undo Activation Steering? Behavioural Recovery Without Weight-Edit Reversal

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24988v1 Announce Type: new Abstract: Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention and offering a way to encode alignment prior to release. However, models are routinely fine-tuned after deployment, and it is unknown whether embedded interventions survive this. We study the stability of embedded steering for refusal suppression and brevity induction across five instruction-tuned models (3B-14B) under non-adversarial SFT and RLHF. Behaviourally, preservation tracks the training data: steering degrades when optimisation pressure contradicts the targeted behaviour and persists otherwise, with refusal ablation losing 64% of its effect on average under SFT. Mechanistically, however, the weight edit survives almost untouched even where behaviour reverts: mean vector recovery is $\rho = 0.004$, and the fine-tuning update along the steering direction is near-orthogonal to its pre-edit weight pattern (mean $\cos\theta = 0.074$). When steered behaviour degrades, fine-tuning does not achieve it by dismantling or reversing the steering mechanism itself. Embedded steering is therefore mechanistically durable but functionally vulnerable, and requires behavioural re-validation after downstream training.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24988v1 Announce Type: new Abstract: Activation steering can be embedded directly into a language model's weights, shaping behaviour without inference-time intervention…
站内正文

待翻译:Unsupervised Post-Training of Foundation Models: A Survey

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24982v1 Announce Type: new Abstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We study Unsupervised Post-Training (UPT): update-bearing adaptation on unlabeled inputs whose learning signal is derived from same-lineage model artifacts rather than an external oracle. We catalog 80 strict UPT methods and organize them by the object that supplies the update signal: a prediction statistic, a sample relation, a self-generated target, or an internal evaluator. Beyond inventory, we show how the choice of internal signal and task structure determines whether post-training improves the model or recursively amplifies error. An orthogonal Input Visibility $\times$ Update Persistence view maps deployment regimes and defines a unified framework for UPT selection and evaluation.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24982v1 Announce Type: new Abstract: Foundation-model post-training usually relies on human labels, preference data, stronger teachers, or executable verifiers. We stud…
站内正文

待翻译:The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24952v1 Announce Type: new Abstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this "dialect tax" across the natural language processing pipeline. Using parallel English dialect corpora that hold meaning fixed while varying surface form, we first confirm that LMs recognize matched Standard American English (SAE) and dialectal texts as semantically equivalent. However, we discover further representational gaps corresponding to downstream performance gaps. Across model families and generations, modern LMs still encode dialectal texts unequally during tokenization, pre-training, post-training, and inference. Strikingly, bypassing traditional subword segmentation via a character-level counterfactual tokenizer removes neither input and output asymmetries nor dialectal accuracy gaps. During pre-training, dialect pairs induce more divergent gradient updates than pairs of entirely unrelated SAE documents, indicating that models find semantically equivalent dialectal content harder to learn from than unrelated SAE documents. During post-training, reward models show contextual, unstable dialect preferences, assigning higher values to isolated AAVE-exclusive tokens than to SAE-exclusive tokens, while full reasoning contexts receive task- and model-dependent dialect penalties. Overall, our findings suggest that the dialect tax is encoded and accumulated not by any one step in isolation, but at every step of the language modeling process.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24952v1 Announce Type: new Abstract: Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the…
站内正文

待翻译:Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24920v1 Announce Type: new Abstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24920v1 Announce Type: new Abstract: This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages fr…
站内正文

待翻译:Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:arXiv:2608.24901v1 Announce Type: new Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding the Resonance direction raises the affective score only partially -- in Qwen by +0.29 (approximately 26% of the natural gap). A direct between-direction contrast confirms the shift is facet-specific in Qwen and Llama (not Gemma); we do not, however, establish a matching human-perceived change. Additive cognitive steering produces no measurable change, but a within-domain control shows the cognitive instrument is too coarse to resolve the differences such steering would produce -- unmeasurable, not a clean null. By contrast, Gemma Recognition ablation lowers the classifier's cognitive score even after adjusting for response length. Detection does not imply reliable control under global interventions, and cognitive-empathy claims warrant an explicit measurement-sensitivity check.

  • AI 服务暂时不可用,系统已先保留来源内容与降级元数据。
  • arXiv:2608.24901v1 Announce Type: new Abstract: A decodable "empathy" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-p…
站内正文

主题导航

研究 — AI 话题新闻 | AI News Hub