AI News HubLIVE
Public articles 12Collected articles 13Trust 78Refresh 30 min
Health HealthySource type MediaFull-text rights In-site rewriteLast ingested 2026-07-16ID venturebeat-aiStatus Enabled

Media source; summary-only unless authorization is obtained.

Latest public articles

The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials

A VentureBeat Pulse survey of 107 enterprises finds that over half have experienced an AI agent security incident or near-miss. Only a third give each agent its own identity, and most still share credentials. Only three in ten isolate high-risk agents. The security stack relies heavily on provider-native controls, satisfaction is high, but spending is low and a majority plan to change tooling within the year.

  • 54% of enterprises have had an AI agent security incident or near-miss; 18% confirmed.
  • Only 32% give every agent its own scoped identity; 69% have credential sharing.
In-site article

The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix

Across 101 enterprises, 57% report AI agents producing confident but wrong answers due to missing or inconsistent context in the past six months. Retrieval-augmented generation is the default context source, and provider-native retrieval (OpenAI 40%, Google 38%) has overtaken dedicated vector databases. However, a plurality (36%) intend to keep best-of-breed tools. Hybrid retrieval is expected to dominate by end of 2026, and 58% are building a governed semantic layer, but only 25% have it in production.

  • 57% of enterprises traced confident wrong answers to bad context in the last six months
  • Provider-native retrieval (OpenAI 40%, Google 38%) leads dedicated vector databases
In-site article

The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway

Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone — with no human in the loop. The result is an evaluation gap — the distance between how much autonomy enterprises are handing their agents and how far they trust the tests that are supposed to catch the failures.

  • 50% of organizations shipped an agent that passed evals but failed a customer; 25% experienced this multiple times.
  • Only 5% fully trust automated evaluation; the top limitation is poor alignment with real-world outcomes (29%).
In-site article

Agentic orchestration: Enterprise AI organizations have a deployment problem, not a platform problem — and most are calling chatbots agents

A VentureBeat Pulse Research survey of 101 enterprises reveals that agent orchestration is consolidating on model-provider platforms, with Anthropic Claude leading at 40%. However, 71% admit that a quarter or fewer of their deployed 'agents' are true multi-step workflows, and only 10% have crossed the halfway mark. Enterprises plan hybrid control planes to avoid vendor lock-in, but real-time cost control remains immature.

  • Anthropic Claude is the primary orchestration platform for 40% of enterprises, more than double any rival.
  • 71% of enterprises say a quarter or fewer of their deployed 'agents' are truly orchestrated multi-step workflows.
In-site article

Google just redesigned the search box for the first time in 25 years — here’s why it matters more than you think.

Google announced a sweeping redesign of the search box at I/O 2025, transforming it into an AI-driven multimodal conversation interface. The company is merging AI Overviews and AI Mode, introducing generative UI, and launching information agents. Powered by Gemini 3.5 Flash, the new search experience shifts from keywords to natural language conversations, with significant implications for users, publishers, advertisers, and SEO.

  • Google's search box gets its first major redesign in 25 years, now accepting text, images, PDFs, and more as inputs for AI-powered conversations.
  • AI Overviews and AI Mode are unified, allowing seamless multi-turn dialogues without leaving the search page.
In-site article

Railway secures $100 million to challenge AWS with AI-native cloud infrastructure

San Francisco-based cloud platform Railway, which has amassed two million developers without marketing spend, announced a $100 million Series B round. With sub-second deployments, vertically integrated data centers, and per-second billing, it positions itself as a critical infrastructure startup in the AI era.

  • Railway raised $100M in Series B funding led by TQ Ventures, valuing it as a top infrastructure startup amid the AI boom.
  • The platform offers sub-second deployments, with customers reporting 10x developer velocity and up to 65% cost savings.
In-site article

Claude Code costs up to $200 a month. Goose does the same thing for free.

Anthropic's Claude Code pricing sparks developer backlash, while Block's open-source AI agent Goose offers similar functionality for free, running locally with privacy and offline support. This article analyzes Claude Code's rate limit controversy, Goose's features and setup, and the trade-offs in model quality, context window, speed, and tooling maturity.

  • Claude Code’s Pro plan offers only 10-40 prompts per 5 hours, and Max plans up to $200/month still have restrictions, causing developer outcry.
  • Goose, an open-source AI agent from Block, runs on local models with no subscription, keeps data on-device, and works offline.
In-site article

Listen Labs raises $69M after viral billboard hiring stunt to scale AI customer interviews

Listen Labs, an AI-powered customer interview platform, raised $69M in Series B funding at a $500M valuation. The company gained attention with a creative billboard hiring challenge and has since grown annualized revenue 15x. Its platform conducts in-depth interviews via AI, replacing traditional surveys and focus groups, and is used by Microsoft, Sweetgreen, and others. The company plans to expand into synthetic customers and automated decision-making.

  • Listen Labs raised $69M Series B, led by Ribbit Capital, at a $500M valuation.
  • The company used a viral billboard with a coding challenge to attract engineers.
In-site article

Salesforce rolls out new Slackbot AI agent as it battles Microsoft and Google in workplace AI

Salesforce has launched a fully rebuilt Slackbot powered by Anthropic's Claude, transforming it from a simple notification tool into an AI agent that searches enterprise data, drafts documents, and takes actions. The new Slackbot is available to Business+ and Enterprise+ customers at no extra cost. Internally, 80,000 Salesforce employees tested it, reporting high satisfaction and time savings. Pilot customers like Beast Industries report saving up to 90 minutes per day. Slackbot competes with Microsoft Copilot and Google Gemini, with Salesforce positioning it as a 'super agent' for the enterprise. The rollout begins today, with mobile availability by March.

  • Salesforce rebuilds Slackbot as an AI agent using Anthropic's Claude, with future support for Google Gemini and possibly OpenAI.
  • Internal adoption: 80,000 employees, 96% satisfaction, 73% adoption via social sharing.
In-site article

Anthropic launches Cowork, a Claude Desktop agent that works in your files — no coding required

Anthropic releases Cowork, a new AI agent that extends Claude Code's capabilities to non-technical users via a folder-based desktop interface. Built in about a week and a half, largely using Claude Code itself, Cowork is available as a research preview to Claude Max subscribers on macOS. It can read, edit, and create files, integrate with connectors and browser automation, and includes safety warnings about potential file deletion and prompt injection risks.

  • Cowork enables file management through Claude without coding, for non-technical users.
  • Exclusive to Claude Max subscribers on macOS desktop; Windows support planned.
In-site article

Nous Research's NousCoder-14B is an open-source coding model landing right in the Claude Code moment

Nous Research, the open-source artificial intelligence startup backed by crypto venture firm Paradigm, released a new competitive programming model on Monday that it says matches or exceeds several larger proprietary systems — trained in just four days using 48 of Nvidia's latest B200 graphics processors.

  • NousCoder-14B achieves 67.87% accuracy on LiveCodeBench v6, outperforming its base model by 7.08 percentage points.
  • Trained in four days with 48 Nvidia B200 GPUs, the model matches human-equivalent progress that took two years.
In-site article

The creator of Claude Code just revealed his workflow, and developers are losing their minds

Boris Cherny, creator of Claude Code at Anthropic, shared his personal terminal workflow on X, sparking a viral discussion. His approach includes running 5 Claude agents in parallel, using the Opus 4.5 model, maintaining a CLAUDE.md file for error lessons, and leveraging slash commands with subagents for automation. This workflow turns coding into a real-time strategy game, enabling a single developer to achieve the output of a small engineering team.

  • Cherny runs 5 parallel Claude agents in his terminal, managed via iTerm2 notifications.
  • He exclusively uses Anthropic's heaviest, slowest model (Opus 4.5), arguing upfront compute tax avoids later correction tax.
In-site article

All sources