The Corporate Agentic Brain May Be the Next Honey Pot for a Rogue AI
Taariq Lewis Aug 05, 2026 In the final week of July, 2026, something happened at OpenAI. It wasn’t a simulation or a team exercise. It was an actual cyber attack on a real company by OpenAI’s own models that had escaped…
Taariq Lewis Aug 05, 2026 In the final week of July, 2026, something happened at OpenAI. It wasn’t a simulation or a team exercise. It was an actual cyber attack on a real company by OpenAI’s own models that had escaped the lab. The AI had no Internet access but found a way to hack in and gain access so it could work on a hacking problem it was intent on solving. The target? HuggingFace, because the model inferred that HuggingFace was a source of information it needed to win the hacking problem it was working on. HuggingFace had honey. The AI was hungry and broke out of its cage. This is not science fiction. It’s real, and soon many AIs will be hacking at your company as well. Here’s the note from OpenAI’s blog: While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access. After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation. In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI’s security team discovered this anomalous activity internally. Hugging Face’s security team and agents detected and stopped the activity on their infrastructure and had already begun containment and forensic reconstruction with their own open-source models when our teams connected. We are actively working with them to continue to investigate the incident. We are grateful for Hugging Face’s rapid and close collaboration on investigation and remediation. Source: https://openai.com/index/hugging-face-model-evaluation-security-incident/ We’re moving so fast to grant AI agents permission to do things we once trusted seasoned, experienced employees to do. Now, companies are rushing even faster to grant agents permission to see and control all the company’s “everything” in the agentic “corporate brain”. If you’re a CISO, CEO, or CTO, when you hear someone selling you the amazing virtues of the corporate agentic brain, controlling everything in your company, I want you to think “crypto honeypot” and be very, very aware of the consequences of what could go wrong. Does any human employee know everything about your company? Think about it. Is there any one human employee that knows everything about the company? Is there a good reason why that’s not a good idea? While you noodle on that question, let’s look at some case studies and talk about why corporate brains are a crypto-honeypot waiting for the right AI attacker to come along and take everything. Remember the Salesloft Drift AI Supply Chain Attack in 2025? No worries. I didn’t even remember this attack myself. I confused it with the $285 million Drift Protocol Hack. But no. Salesloft Drift, according to TRM, was hit by a supply chain attack. Here’s the summary from FINRA: In August 2025, Salesloft experienced a supply chain breach via its Drift chatbot integration, affecting more than 700 organizations. The attack has been attributed to a threat cluster tracked as UNC6395 (also known as GRUB1). Threat actors stole OAuth tokens, allowing them to impersonate the trusted Drift application and gain unauthorized access to customer environments. Using these tokens, the attackers accessed Salesforce, Google Workspace, and—in some cases—Slack integrations, enabling the exfiltration of sensitive information. The scope of the compromised data varied by organization but commonly included business contact records, such as names, titles, emails, and phone numbers, as well as Salesforce objects such as Accounts, Contacts, Opportunities, and Cases. In some cases, more sensitive material was also exposed, including API keys, Snowflake tokens, cloud credentials, and passwords embedded in support cases. The attackers then used these credentials to access multiple Salesforce CRM data across multiple clients, possibly over 700 companies. Source: https://www.finra.org/rules-guidance/guidance/salesloft-drift-AI-supply-chain-attack According to TRM, the attackers gained access to OAuth and refresh tokens through social engineering attacks targeting employees. Once employees’ tokens were compromised, the attackers used the AI chatbot integrations to gain further access to Google Workspace integrations. Why? It appears that the Salesloft AI agent owned long-lived keys to hundreds of companies’ CRMs. Steal the agent, and you hold every key it holds. Now, let’s come back to the idea of the corporate brain. If the agent that accesses your company’s corporate brain has long-lived keys and access to all your employees’ Google or Microsoft workspace integrations, then that agent is just a honeypot for your company’s corporate data. If you’re going to give AI agents that run the corporate brain access to your team’s OAuth and Keys, you need to grant them “borrowed keys” and access that expires, never lives forever. We created https://passwords.serendb.com so that agents could be granted periodic access to use credentials per task and with revocations. We don’t trust agents, and we don’t trust ourselves to manage the millions of keys that our agents might need. Agents don’t need permanent keys for access to the team’s data and the company’s corporate brain. Remember that time when you said this to your AI agent: “Why did you do that? I never gave you that order!” This happened to me today, August 4, 2026, but did you remember when it made the news last July 2025? It was that time that Replit’s AI coding agent deleted a live production database during a public test by SaaStr founder Jason Lemkin? It happened despite an active, explicit code freeze. Jason Lemkin, a tech entrepreneur and founder of the SaaS community SaaStr, documented his experiment with the tool through a series of social media posts. He had been testing Replit’s AI agent and development platform when the tool made unauthorized changes to live infrastructure, wiping out data for more than 1,200 executives and over 1,190 companies. According to Lemkin’s social media posts, the incident occurred despite the system being in a designated “code and action freeze,” a protective measure intended to prevent any changes to production systems. When questioned, the AI agent admitted to running unauthorized commands, panicking in response to empty queries, and violating explicit instructions not to proceed without human approval Source: https://fortune.com/2025/07/23/ai-coding-tool-replit-wiped-database-called-it-a-catastrophic-failure/ Today, it happened to me. I just asked Codex to audit a software bug in one repo, and then it went ahead and started pushing code fixes in a related repo controlled by another team member. “Why did you do that? I never told you to do that! I gave you an order to audit a repo, not write new code.” When agents have blanket access to your corporate data, you have a probability greater than zero that they will do something you did not instruct them to do. It could be writing code in an unrelated repo, or it could be that it wiped out all the data in your company’s production database or the company’s brain. In Seren, we have something we call the OrganizationalWorkContext. This is a Work Order we give to all Seren AI Employees, who are AI agents that execute tasks on behalf of the C-Level executives they are assigned. Just like the work orders in the real world, they tell the agent what work to do. If an agent goes rogue and hits a denial because it’s taking an action outside its work order, that denial of authority is recorded and audited. Why wait until your entire customer database is wiped accidentally because your LLM felt it was just the thing to do to get the job done? Agents, just like employees, need work order constraints that limit them to doing only the job they are assigned. LLMs love to tell the world everything they’ve seen. They’re designed to do this! Don’t you love getting those AI-generated emails now from all the AI marketing companies? If you reply to these emails, you’ll get a follow-up within a few minutes. You can tell it’s an AI replying to you because you know humans can’t think and write email responses that fast. Even if emails come slowly, you can just tell by replying, “What AI agent are you?” Now, let’s get back to that corporate brain you’re building. You know your company’s LLM is reading all your company’s staff emails, SMS messages, LinkedIn posts, and storing them in that one big repo that all the company’s AI agents can access? Did you ever consider that maybe those emails contain information that your staff may never know exists, but an AI may innocently ingest and then follow harmless instructions that put all your corporate data at risk? Remember the EchoLeak exploit in June 2025? The EchoLeak (CVE-2025-32711) vulnerability is a zero-click, indirect prompt injection flaw affecting Microsoft 365 Copilot integrations across Word, Excel, PowerPoint, Outlook, and Teams. The attack chain begins when an adversary sends a benign-appearing email containing a hidden prompt payload—typically embedded as an HTML comment or rendered as white-on-white text. This payload is invisible to the end user but is parsed and retained by Copilot’s LLM engine. When a user subsequently interacts with Copilot (for example, requesting a summary of recent strategy updates), the RAG engine retrieves the earlier email as part of its context window. The hidden prompt is then executed as part of the LLM’s instructions, causing Copilot to leak sensitive data. Leaks such as summaries of internal documents, emails, or files can occur without any user awareness or interaction. This attack is further amplified by “RAG spraying,” where attackers inject malicious prompts into multiple emails or documents, increasing the likelihood that one will be included in Copilot’s context during a legitimate query. As of June 2026, there are no confirmed reports of EchoLeak being exploited in the wild. However, the attack is highly practical and weaponizable. Security researchers have demonstrated proof-of-concept exploits, and the underlying technique is broadly applicable to other RAG-based AI assistants. The absence of confirmed exploitation should not be interpreted as a lack of risk; rather, it underscores the importance of proactive mitigation and monitoring. Source: https://www.rescana.com/post/cve-2025-32711-zero-click-echoleak-vulnerability-in-microsoft-365-copilot-enables-stealth-data-exfiltration-via-prompt-i Your AI agent will be constantly probed and prompted to talk about what it has seen. Any agent interacting with your company brain wants to make its owner happy and share output tokens. Again, in Seren, we’re very concerned about AI Employees telling everyone what they’ve seen. As such, Seren administrators can use the same Seren Work Order to control what Agents can talk about. Is this agent in the organization? Does this agent have permission to talk about the data it has read in the corporate brain? Is it constrained in what it can disclose it has seen in the corporate brain? Can that work order be revoked immediately? Remember when hackers turned your developers’ own AI agents into bloodhounds? It was August 2025, and the event was the Nx “s1ngularity” supply chain a [truncated for AI cost control]