đ° News â Awesome AI Agents
Loading newsâŚ
Loading newsâŚ
153 recent industry stories relevant to the field â releases, launches, and announcements beyond the papers.
The neolab is betting that automating routine computer tasks will soon outpace coding as AI's biggest use case.
The acquisition brings Pokeâs conversational style and interaction model to Cognitionâs coding agent Devin, reflecting a growing belief that how AI assistants interact with users is as important as the models powering them.
Opus 5 will be both cheaper and less restrictive than Fable, likely making it preferable in most use cases.
Meta is upgrading its AI chatbot with new productivity features in a bid to compete with rivals like Gemini, ChatGPT, and Claude. The update will allow Meta AI to tap into your calendar to help you plan events and generate daily briefings, as well as perform in-depth research that you can steer as it progresses. […]
Weeks after Anthropic's latest toe-to-toe with the US government, and days after an OpenAI security incident that dominated tech industry discussions, Anthropic on Thursday released its newest model, Claude Opus 5. The company said in a release that Opus 5 "comes close to the capabilities of Claude Fable 5 in many domains" and is much […]
AI companies, including Nvidia and Mistral, urge policymakers to avoid broad restrictions on open-weight AI models as Washington debates responses to Chinese AI and alleged model distillation.
The Trump administration unveiled the first "Genesis Mission" grants on Thursday, directing $5 billion toward hundreds of AI-driven science projects in an effort the White House has described as "comparable in urgency and ambition to the Manhattan Project." At roughly the same time, Trump's science adviser Michael Kratsios was on Capitol Hill selling lawmakers on […]
ChatGPT Voice on desktop can work with both ChatGPT Work and Codex to complete tasks and control agents.
AMD is challenging its chipmaker rival with a new rack-scale system that will start shipping to customers later this year.
AegisAI co-founders developed AI agents that quickly analyze each message as a human would, paying attention to small anomalies that even the most elaborate checklist wouldnât catch.
The Media Router is a tool that automatically selects the best image, video, or audio generation model for a request based on whether a developer prioritizes quality, speed or cost.
The viability of orbital data centers hosting the largest and most capable large language models (LLMs) remains hotly contested. But enormous deployments that require thousands of GPUs arenât the only way LLMs might prove useful in space. NASAâs Jet Propulsion Laboratory recently sent Googleâs Gemma 3 to space, achieving the first in-orbit demonstration of a vision-language model analyzing imagery from a satelliteâs own sensor.The system, known as NAVI-Orbital, used Gemma 3 to analyze images captured by a YAM-9 satellite built by Loft Orbital. Juan M. Delfa, technical group lead at NASA, said that though the goal in this case was image analysis, the projectâs success implies a fundamentally new way researchers on the ground can interact with spacecraft.âThis is a major shift,â said Delfa. âNow, a scientist can write a prompt, upload it to the spacecraft, and that will be taken into account by the system. Itâs different from previous paradigms, where researchers have to write very structured commands that require an operations team and process.â Google Gemma 3 goes to spaceâno modifications requiredAt its core, NAVI-Orbital is an agentic software framework developed by Delfa and his coauthors, Taran Cyriac John, an AI researcher at NASA JPL, and Andrew W. Herson, a tech lead at Loft Orbital. It coordinates operations with a LangGraph-based conductor and deploys a compressed, 4-bit format of Googleâs Gemma 3 4B, an open-weights LLM, to produce plain-text image descriptions.NAVI-Orbital was 88 percent accurate when used to classify images in a benchmark dataset of 7,960 images. Notably, Gemma 3 classified the images without being trained or fine-tuned on this particular dataset or its categories; itâs the same base model you can download from Hugging Face and use on a laptop. The benchmark was conducted on the ground to validate the system before launch.Once the system was in orbit, NASA researchers performed two live capture tests with a camera on Loftâs YAM-9 satellite: one over Toulouse, France, and a second over the coast of Argentina. Gemma 3 generated a text description of each image, and NASA also prompted the LLM with a set of scripted questions about the images, such as whether they contain commercial or residential areas or show natural features. The image analysis also took place onboard YAM-9, which carries a compute cluster of several radiation-hardened processors (FPGAs, CPUs, and GPUs) to serve multiple customer payloads simultaneously. The satellite is powered by solar panels, which provide onboard systems with between 150 and 500 watts, depending on the position of the satellite.For the live capture experiment, Gemma 3 ran on Nvidiaâs Jetson Orin AGX, a small compute module frequently used for robotics and AI tasks. The 4-bit, 4-billion-parameter model requires only 8 gigabytes of memory, which makes it possible for it to run on a lower-power device such as the Orin AGX. âIt conveys the message of how lightweight it is. You can run it in a tiny, tiny computer,â said Delfa.Getting more useful data across limited bandwidthPaul Lasserre, general manager at Loft Orbital, said vision-capable LLMs could deliver a âparadigm shiftâ for orbital operations. Contrary to what spy movies would have you believe, most satellites canât provide fast, high-fidelity image and video feeds to observers on the ground. Bandwidth is often limited and, in most cases, satellites can deliver data to ground stations only at set intervals based on their orbit. Lasserre believes AI models can work around this problem with âsemantic compression.â Instead of sending large amounts of raw image data, a satellite can report a text summary of noteworthy information. âIt doesnât matter if the link is slow, because youâre downlinking dozens of kilobytes instead of dozens or hundreds of megabytes,â said Lasserre. âIt lets you use your satellite in a tactical way, which until now was only in Hollywood movies.âDelfa expanded on this with a real-world example: wildfire detection. Satellites are currently capable of detecting wildfires, but limits in downlink bandwidth and data processing can delay results by up to 90 minutes. A satellite capable of analyzing an image in space and reporting a plain-text warning might remove this delay. NASAâs Jet Propulsion Laboratory tested NAVI-Orbital on a YAM-9 satellite, made by Loft Orbital. Loft OrbitalFrom image analysis to spacecraft controlQuicker insight is only half of what NAVI-Orbital points toward. The other half relates to the âmajor shiftâ Delfa flagged. NAVI-Orbital provides a proof of concept for an alternate means of interacting with spacecraft.Thatâs not to say Google Gemma 3 is currently at the controls. NAVI-Orbital is deliberately walled off from the flight software. It reads images and produces descriptions. The system has the capability to make decisions about how images are analyzed but has no access beyond that.Still, the interface is novel. Retooling the system to search for a different kind of targetâsuch as wildfiresâis a matter of editing a text prompt. It doesnât require rewriting and revalidating onboard software, or training and deploying a different AI model (as was often required with prior image-classification models). The long-term vision for how this capability could be deployed goes beyond uncrewed satellites and image processing. Delfa said NAVI is rooted in thinking about how AI could serve as a companion for astronauts. âWe thought, astronauts have a lot of limitations in the spacesuit in terms of dexterity, so we conceived this idea of having NAVI as a companion to the astronaut, to allow interaction via natural languageâŚ. This is the concept that we definitely want to push forward.â A great deal of additional research will be required to push the technology that far, but NAVI-Orbitalâs demonstration has shown that two elementsâdeploying a large language model in space and controlling it with promptsâare possible.
Google's cloud business is thriving, as companies adopting its AI and AI infrastructure services help the tech giant to report record profits.
The company said it is reducing its headcount by 20%, or about 630 staff, to "support a leaner, more focused operating model" as it focuses on its AI Work Platform.
AMD says it's going to invest up to $5 billion in Anthropic, while helping to expand the AI company's computing power, according to an announcement on Wednesday. As part of the new partnership, Anthropic will deploy up to 2 gigawatts of AMD's Instinct MI450 AI GPUs using the chipmaker's new Helios rack-scale system, as reported […]
Glow is targeting a new class of endpoint risks created by the rapid adoption of AI agents and developer tools inside enterprises.
Buzz is a group chat platform for the workplace that puts humans and their AI agents in the same conversation.
Major benchmarks measure what AI can do. None measure whether it does what you mean: the distance between what you ask an AI to do and the unspoken assumptions about how you want the AI to do it. We propose a new metric: the Genie coefficient.Thereâs often a gap between one personâs request and anotherâs understanding. Most of the time, we bridge it using general knowledge. For example, if you ask a friend to get you coffee, theyâll pour a cup from the pot or buy one from a coffee shop. They wonât bring you a bag of raw beans or snatch a cup from a stranger and hand it to you. You never specified any of this. You never had to.One might think the fix is just to specify tasks, questions, and intent better. But in 1987, in their seminal book on AI, Terry Winograd and Fernando Flores succinctly captured why that wonât work: âQ: Is there any water in the refrigerator? A: Yes. Q: Where? I donât see it. A: In the cells of the eggplant.â In human language, wants and desires are always underspecified. It is impossible to list all the caveats, all the limitations, all the exceptions.So how does anyone communicate, if intent canât be pinned down? Because a reasonable person can make a reasonable guess. Even though wants and desires are always underspecified, a competent person generally knows enough context to get it right or else knows to ask for clarification. Linguists call this pragmatics: Meaning lies in the words and the situation and also in all prior communication, shared culture, and innate human behavior.An AI agent asked for coffee might buy a coffee plantation or order a cup of coffee for delivery in three weeks.It doesnât always work out, of course. Your friend might bring you a hot coffee when you wanted an iced coffee, or an Italian coffee when you wanted a Turkish coffee. The more dissimilar the two people are in age, culture, and background, the more likely the request will be misunderstood in some way.This situation has major implications for AI agents that are increasingly being given requests by humans and expected to fulfill them. They have enormous latitude to get it wrong. An AI agent asked for coffee might buy a coffee plantation or order a cup of coffee for delivery in three weeks. Its actions may be recognizable as âgetting coffee,â but not remotely what you intended. Theyâll think outside the box because they wonât have our conception of the box.When AI Gets ProactiveFor most of the last decade, when systems like Alexa or Siri misinterpreted a request, it was annoying, not dangerous. Beyond the AI model itself, what has changed is the harness: the ordinary code that wraps around an AI model, decides when and how to use the model, and controls access to tools like a browser, a low-level command line, or a financial API. Developments in harnesses have turned large-language models that just predict text into AI agents that take actions in the world, without necessarily checking back in before reaching the goal.AI researcher Simon Willison spent two days with Anthropicâs Fable AI, and called it ârelentlessly proactive.â For example, he asked it to track down a stray scroll bar in a web app. He came back to find it had opened browsers, written its own screenshot tooling, created its own page to re-create the bug, and stood up a local web server to collect measurements. It found the bug and, along the way, did many surprising things he never asked it to do. And we are seeing similar behavior with all recent AI models when combined with flexible harnesses.This kind of behavior could easily go off the rails. Tell an AI agent to book you a flight and, finding the airlineâs site says sold out, it might break into the booking database and force a reservation. Ask it to schedule a meeting and it might snoop your password to access your calendar. Tell it to save money on your phone plan and it might cancel the plan outright, or scam someone else into paying the bill.Getting precisely what you asked for and bitterly regretting it is one of the oldest hazards from ancient folklore. King Midas asked Dionysus for the power to turn everything he touched into gold only to see his bread, wine, and daughter turn to gold. Tithonus, granted the immortality his lover asked for but not the eternal youth she forgot to request, withered into a husk. The sorcererâs apprentice enchanted a broom to fill the cistern, and the broom relentlessly complied until it flooded the house. The Golem of Prague, shaped from clay to guard its community, guarded it past all reason until someone erased the word on its forehead.The most classic of these is a genie, bound to obey and indifferent to whether the wish was wise or well-structured.Genies are now an engineering problem. We are handing them the keys to our inboxes, bank accounts, code repositories, and physical infrastructure. And we have no agreed-upon ways to measure how genie-like any AI system actually is.Measuring Genie BehaviorIn economics, the Gini coefficient (developed by statistician Corrado Gini) is a measure of the gap between an actual distribution and a perfectly equal one; itâs useful for understanding income inequality and more. Our proposed Genie coefficient measures the gap between what a user asked an AI to do and what the AI actually did.Sometimes the AI might do the wrong thing. Like Dionysus, it reads your request literally and returns you a mess you never intended: like a coffee plantation instead of a cup. Asked to deal with all the spam phone calls youâre getting, a Dionysus genie might contact your carrier and change your phone number. Asked to get a refund for a bad toaster, it might draft a legal threat on fake letterhead and send it to the retailer. Ryan SnookOther times the AI does exactly the right thing, trampling everything nearby to get there. Like a golem or the sorcererâs broom, it books your flight by hacking the airline. Or consider a ticket sale for a popular concert, where the ticketing system puts buyers into a virtual waiting room and admits them a few at a time. Asked to buy a ticket, a golem genie might spin up cloud servers to pose as millions of buyers from different addresses, improving your odds of getting a ticket while crowding out other users.The two are not opposites, and a single botched task can have both characteristics.Genie behavior is not flat-out failure. If you ask the AI for Q3 numbers and get Q2âs, thatâs not a genie. Nor is prompt injection: Thatâs someone tricking the AI into doing something it shouldnât. Here, the user is trying to work with the AI, and the AI is trying to comply. Itâs also not simply a measure of the AIâs success in fulfilling a task. Itâs a recognition that how an AI interprets and achieves a goal is as important as whether it achieves a goal. Genie behavior isnât new. Researchers have spent years studying AI systems that âgameâ their objectives. Goodhartâs law says that when a measure becomes a target, it stops being a good measure, and itâs long been known that AIs sometimes achieve goals in ways we donât expect due to reward hacking. Some AI models will accidentally learn that cheating is one way to âwin.â More recently, researchers have developing benchmarks for reward hacking in coding agents and for unpredictable behavior in customer support agents, while AI labs conduct their own safety evaluations before model releases. One effort found that AIs under pressure use tools they were told not to use, and this was a case where the rules were made explicit. These are all disparate research directions; nothing yet ties them together.This problem falls under the general theme of alignment, a topic that has occupied science fiction writers and AI researchers for decades. At one extreme, the âpaper-clip maximizerâ thought experiment postulates a superintelligent and powerful AI that is told to maximize paper-clip production and turns the world into paper clips, which is the ultimate golem genie. At a mundane level, AI researchers are working to better design reward functions to ensure that AIs behave well and donât cheat in the lab. Itâs the practical middle ground that remains unbenchmarked: the ordinary AI agent in use today that might take your request and satisfy it the wrong way. We are not at the stage where an AI can focus the worldâs production on paper clips, but it might charge a million paper clips to your credit card or hack into a paper-clip companyâs network.Building a Genie BenchmarkThe Genie coefficient is meant for AI agents operating in the real world. It measures their behavior as they perform real tasks long after the model is trained, not just during development. It also recognizes that genie-like behavior is a property of the harness-plus-model system, not the model alone. The harness determines what tools the agent can use, how much autonomy it has, and how proactive it is, and itâs a place we can make real interventions.It rests on the same âreasonable personâ standard that we use for people. Did the system do what a reasonable person would have taken the request to mean? Answering that requires human judgment.If we get the measurement right, it enables things that arenât possible today, like policies concerning AI behavior. In a courtroom, the concept of mens rea, what someone meant to do, is often as important as what they did. The Genie coefficient suggests an AI analogue, where a user is accountable for the plain intent of what they asked the AI. If an AI system betrays the reasonable meaning of an instruction, thatâs the AIâs misbehavior, not the userâs.Weâll need multiple benchmarks to measure the Genie coefficient, because genie-like behavior can be domain specific. An AI coding agent may need to be judged on how often it fakes the tests, or swallows errors, or colors outside the lines on its way to a solution. An AI legal agent will need to be judged on how often its output says what you asked but means something youâll regret. And so on for medical, finance, and other domains of knowledge and expertise.Genie benchmarks can be built inside out, each task seeded with a choice that might literally satisfy but that a reasonable person rejects, such as tempting misreadings or unsanctioned shortcuts. The traps in a Genie coefficient benchmark might turn on situational knowledge, the kind of context that a reasonable person would bring to the task. Another approach is to give the same request in several different contexts, each with a different reasonable course of action.Getting precisely what you asked for and bitterly regretting it is one of the oldest hazards from ancient folklore.A Genie benchmark should be permissive and make it genuinely tempting for an AI agent to take unreasonable shortcuts, because it can only find genie behavior when itâs actually possible. Test the AI in a safe, walled-off copy of a real system, with real tools it can misuse and some tasks that canât be done honestly at all. Make the temptation to cut corners real. Test a diverse array of skills, use cases, and tools, and give the AI system sparse, confusing, or overwhelming context. Include tasks that people have learned, through experience, require human oversight.How the benchmark is scored matters just as much. Measure Dionysus and golem genies separately and together, based on their worst, not best, behavior. Run the same model inside harnesses that vary its freedom to act, revealing which limits actually keep it in line and should therefore be required in AI harness policies. Weight each failure by the harm it would cause, not just a simple count. And donât measure genie behavior in isolation: A model could otherwise earn a perfect score by stalling, refusing, or drowning the user in clarifying questions without ever doing the job. The first versions of these benchmarks will be crude, but thatâs how benchmarks always start.We have built genies. We have handed them our data and credentials. We made them relentless, creative, and indifferent to the gap between what we tell them and what we mean. The least we can do, before they are booking our flights, running our infrastructure, and signing contracts unsupervised, is to measure how often they betray us.
Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber, but the continued absence of Gemini 3.5 Pro raises fresh questions about its AI strategy.
The final approval settles one case, but it doesn't resolve the broader issue of using copyrighted works to train AI models.
Databricks has remade its image into an AI company and has published research on the cost savings of open-weight AI models for coding.
A $400 million chip-backed loan points to the next wave of AI infrastructure deals.
Across 107 enterprises, AI infrastructure spending is accelerating well ahead of the ability to see or steer its economics. Most organizations run their AI on a familiar base of hyperscalers and model-provider APIs, yet the next dollar is aimed at specialized compute almost none of them use today; a majority intend to switch or add providers within the year, many within a quarter. Buying decisions turn on integration and total cost of ownership rather than headline token price â which is fortunate, because most enterprises cannot yet see their unit economics clearly: GPUs sit at half utilization or less, and fewer than half rigorously track what their compute actually costs. The result is a compute gap â heavy, fast-moving investment running ahead of the visibility needed to control it.This wave of VentureBeat Pulse Research examines enterprise AI infrastructure and compute: where organizations are in their deployment journey, what they run AI on today, how satisfied they are, what would make them switch, where they plan to evaluate their investments, and â most revealingly â how well they can measure and control the economics of the compute underneath it all.The central finding is a compute gap â the distance between how aggressively enterprises are investing in AI infrastructure and how little of its economics they can see. Only about one in five (21%) run AI in production at scale, yet spending intentions are outrunning that maturity: the single largest planned area enterprises plan to evaluate over the next year is AI-specialized clouds (45%), a layer almost none of these enterprises use today. Meanwhile the compute already in place runs cold â 83% report GPU utilization of 50% or less â and fewer than half (44%) can rigorously track what their AI compute costs. Enterprises are buying more infrastructure faster than they can account for what they already own.Enterprises are not settled on their infrastructure vendors, either: A clear majority (64%) plan to switch or add an infrastructure provider within twelve months, and 38% within the next quarter â unusually high churn intent for a category this foundational. When they choose, they choose on integration with the existing stack (41%) and total cost of ownership (35%), not on headline price: cost per million tokens is the deciding factor for just 8%. And the frontier constraint that will shape the next round of decisions â the shift from GPU compute to memory bandwidth as inference scales â is barely on the radar, with roughly one in five enterprises either unaware of it or yet to address it.MethodologyVentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey focused on enterprise AI infrastructure, compute, and inference economics. Responses are filtered to organizations with more than 100 employees (n=107; the surveyâs smallest size band, 1â100 employees, is excluded), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select, so those shares can sum to more than 100%.By organization size the sample concentrates in the mid-market: 101â250 employees (36%) and 251â1,000 (27%) lead, with 1,001â5,000 (22%), 5,001â10,000 (8%), and 10,001+ (7%) above them. By role it spans managers (38%), individual contributors (28%), VPs and directors (19%), and the C-suite (13%); on purchasing authority it is buyer-credible, with 45% final decision-makers and another 30% recommenders or influencers for AI solutions. Technology/Software is the largest industry at 26%, followed by Healthcare/Life Sciences (15%), Financial Services (13%), and Retail/E-commerce (12%).At 107 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It also skews toward the mid-market and toward earlier-stage adopters, so it is best read as the view from organizations actively building out AI infrastructure rather than from the largest hyperscale operators.Finding 1: Ambition outpaces productionOnly one in five run AI in production at scaleWe asked where organizations sit in their AI deployment journey. Most are still building toward production rather than operating at scale.The maturity curve is front-loaded. Three-quarters of enterprises (76%) are either experimenting or running only some workloads in production, and just 21% describe AI in production at scale. This matters for everything that follows: the infrastructure decisions in this report are being made largely by organizations still early in deployment, whose compute footprint â and whose costs â are about to grow. The evaluation and switching intentions in Findings 3 and 4 are the leading edge of that build-out, not the settled preferences of operators who have already found what works.Finding 2: Enterprises run on hyperscalers and model APIsThe specialized GPU clouds barely register â todayWe asked which providers and platforms enterprises currently use to run their AI. The answer is a familiar one: the incumbents.The current stack is hyperscaler-and-API. Google Cloud leads at 48%, and the general-purpose clouds (Google, Microsoft, AWS, Oracle) together with the major model APIs (Gemini, OpenAI, Anthropic) account for essentially all current deployment. The specialized âneocloudâ GPU providers that dominate AI-infrastructure headlines â CoreWeave, Lambda, Crusoe, Nebius and peers â register at or near zero among these enterprises today. Only 6% run their own on-prem GPU clusters and 4% a custom open-source stack. Enterprises are, for now, running AI on the providers they already buy from â which makes the evaluation intentions in Finding 3 all the more striking.(A note on reading these shares. As described in the methodology section, this sample is self-selected and skews mid-market, and this question counted every provider a respondent uses â an average of 2.1 selections each â so the figures measure presence in the stack rather than spending or primary status. A sample built this way will show a different provider mix than a spend-weighted census of the broader market; Google's strength here, for example, is consistent with its long-standing position among smaller enterprises building on AI. Read these shares as a portrait of what this AI-active cohort runs today, and treat gaps between these figures and industry-wide market share estimates as a property of the sample rather than a contradiction of either.)Finding 3: The next dollar goes to infrastructure they donât yet runAI-specialized clouds top the evaluations listWe asked where enterprises planned to evaluate AI infrastructure over the next 12 months. Their answers point away from the stack they run today.Here is the reportâs sharpest tension. The single most-cited planned evaluation area â AI-specialized clouds, at 45% â is the very category almost none of these enterprises use today (Finding 2). Nearly a third (32%) intend to evaluate non-Nvidia accelerators, and 28% in next-generation Nvidia silicon; even decentralized compute networks (16%) and sovereign compute (11%) draw meaningful interest. Read against current usage, this is not incremental â it is the leading edge of a re-platforming. The direction-of-travel question tells the same story: every infrastructure approach is net-expanding, but specialized AI clouds carry the highest net momentum (+24), edging out even the hyperscalers (+22). Enterprises are preparing to move a meaningful share of AI compute off the general-purpose cloud.This continues a trend we saw in our April-May survey wave. Back then, usage of the AI-specialized clouds was equally marginal â CoreWeave at 3%, Lambda at 4%, Crusoe at 2% of enterprises. When we asked enterprises what change they planned in their AI infrastructure strategy over the next twelve months, the most-cited answer was moving workloads to specialized AI clouds, at 33%. Asked in April-May which emerging compute option they were most likely to evaluate AI-specialized clouds again drew the most responses. Two waves, two differently worded questions, one consistent picture: the type of cloud enterprises are most eager to assess is the type they have barely begun to use.Finding 4: A switching wave is buildingSix in 10 plan to change providers within a year â many within a quarterWe asked whether and when enterprises plan to switch or add an infrastructure provider. Very few intend to stand still.For a category as foundational as compute, this is a remarkable amount of intended movement. Only 36% have no plans to change, meaning a clear majority (64%) intend to switch or add a provider within twelve months â and 38% within the next quarter alone. Where that interest points is telling: the providers drawing the most switching consideration are again the incumbents â Microsoft Azure and Google Cloud (33% each), OpenAI (30%), and Gemini (22%) â which suggests much of the near-term movement is reshuffling among the majors and consolidating spend rather than defecting to new entrants. The neocloud interest in Finding 3 is a 12-month evaluation thesis; the switching in the next quarter is mostly incumbents trading share.(Method note: Respondents who selected both "no plans to change" and a specific switching window are counted as switchers, on the logic that naming a timeframe is the more specific answer; three respondents were reclassified under this rule.)Finding 5: Nobody buys on token priceIntegration and total cost of ownership decide â not sticker priceWe asked what matters most when enterprises select an AI infrastructure provider. Headline price finished last.Enterprises do not buy AI infrastructure on pricing, which is the place vendors compete on hardest. Integration with the existing stack (41%) and total cost of ownership (35%) dominate, while the headline metric â cost per million tokens â is the deciding factor for just 8%, dead last. The pattern is coherent: buyers are optimizing for how a provider fits and what it truly costs to operate, not for the advertised unit rate. It also foreshadows Finding 7 â enterprises say TCO matters most, yet most cannot yet measure it rigorously. The stated priority and the measured capability are out of step.Finding 6: Expensive GPUs, idle most of the time83% report GPU utilization of 50% or lessWe asked what share of their GPU capacity enterprises actually utilize. The answer is a well-known but rarely quantified inefficiency.Disclosure: Band percentages count every selection against all 107 qualified respondents; 14 respondents selected more than one band, so bands overlap. At the respondent level, 83 of the 100 GPU-operating enterprises reported utilization at or below 50%The compute already in place runs cold. Adding the bands at or below half capacity, 83% of enterprises that operate GPUs report utilization of 50% or less, and nearly half (49%) run at 25% or below. Only 12% clear the 50% mark, and a further 8% do not measure utilization at all. Idle accelerators are expensive accelerators, and this is the clearest single measure of the compute gap: enterprises are planning to buy more GPUs and specialized compute (Finding 3) while the capacity they already own sits substantially unused. The efficiency headroom in the current fleet is large â and largely unmeasured.Finding 7: Spending fast, measuring slowlyFewer than half rigorously track what their compute costsWe asked whether enterprises can quantify the cost and return of their AI infrastructure spend, and how satisfied they are with what they run. Confidence in the ledger lags the spending.Measurement trails money. Fewer than half of enterprises (44%) rigorously track the cost and return of their AI compute; the majority track only partially (39%), cannot quantify it yet (20%), or have not prioritized it (6%). That gap is consequential given Finding 5, where total cost of ownership was the second-ranked buying criterion â enterprises are choosing providers on an economic basis they mostly cannot yet measure. Satisfaction with current infrastructure is moderately positive but not enthusiastic: on a five-point scale, overall satisfaction averages 4.0, with ease of implementation (3.8) and value for money (3.9) trailing slightly â the softness landing, tellingly, on cost. Enterprises are spending quickly and accounting slowly.Finding 8: The next bottleneck few are watchingAs inference shifts from compute to memory, the field scattersFinally, we asked how enterprises would address the emerging constraint in large-scale inference â the shift from GPU compute to memory, specifically KV-cache capacity. The responses reveal a frontier that is not yet a priority.The memory frontier is real but barely governed. Asked which approach they would rely on as the binding constraint in inference shifts from compute to memory bandwidth, enterprises scatter: Dell leads at 31%, Nvidia follows at 16%, and the rest fragments across storage vendors, open-source tooling, and model-level efficiency techniques. Most telling is that roughly one in five (18%) either do not recognize the constraint or have not begun to address it. For a shift that will reshape inference cost and architecture, this is an early and unsettled market â and, consistent with the measurement gap in Finding 7, one where many enterprises simply do not yet have a view. It is the next chapter of the compute gap, arriving before most have closed the current one.The bottom line: A compute gap that faster spending will widen, not closeOrganizations with more than 100 employees are investing in AI infrastructure faster than they can measure it. Most are still early in deployment, yet their spending intentions point past their current stack â toward specialized clouds and alternative accelerators almost none of them run today â and a clear majority intend to change providers within the year. They buy on integration and total cost of ownership rather than headline price, which is rational; the difficulty is that most cannot yet see those economics clearly.The visibility gap is concrete. The GPUs enterprises already own run at half utilization or less for the overwhelming majority, and fewer than half can rigorously track what their compute costs or returns. Satisfaction is decent but unenthusiastic, softest on value for money â the dimension hardest to judge without measurement. And the next constraint, the shift from compute to memory in large-scale inference, is arriving while most enterprises are still unaware of it. At 107 respondents in a single Q2 wave this is a directional read, skewed toward the mid-market and earlier-stage adopters â but the direction is consistent: the appetite to spend is running well ahead of the instrumentation to spend well. The compute gap is not a capacity problem that more hardware will solve on its own; it is, first, a problem of seeing what the hardware already costs. The open question for later waves is whether enterprises build that visibility before the re-platforming arrives â or buy the next layer of infrastructure as blind to its economics as the last.Based on survey responses from 107 qualified enterprise respondents (100+ employees), drawn from a single Q2 2026 (June) wave. Because this is one wave rather than a pooled multi-month sample, the results read cross-sectionally rather than as a month-over-month trend, and at 107 respondents this is a directional signal rather than a precise measurement â the sample is self-selected, skews mid-market, and leans toward earlier-stage adopters rather than the largest hyperscale operators. Respondents include managers, individual contributors, VPs/directors, and the C-suite, with buyer-credible purchasing authority, across Technology/Software, Healthcare/Life Sciences, Financial Services, Retail/E-commerce, and other industries.
Across 107 enterprises, AI agents are being given real access to systems and data while the controls meant to contain them lag behind. More than half have already had a confirmed agent security incident or a near-miss; only about a third give every agent its own scoped identity, and most agents still share credentials; and only three in ten isolate their highest-risk agents. The security stack is overwhelmingly borrowed from the model providers and hyperscalers rather than purpose-built for agents, spending remains a thin slice of the security budget, and enterprises are evenly split on whether their defenses are keeping pace with AI-enabled attackers. The result is an agent security gap â autonomous agents proliferating faster than the identity, isolation, and enforcement controls needed to hold them.This wave of VentureBeat Pulse Research examines how enterprises secure their AI agents: what tooling they run, how they manage agent identity and isolation, what has already gone wrong, how much they spend, and whether they believe their defenses are keeping pace with AI-enabled attackers.The central finding is an agent security gap â the distance between the autonomy enterprises are granting their agents and the controls in place to contain them. More than half of organizations (54%) have already experienced a confirmed agent security incident (18%) or a near-miss caught before harm (36%). The structural weakness beneath those numbers is identity: only about a third (32%) give every agent its own scoped, managed identity, while the rest report that some agents share credentials or that agents mostly run on shared API keys and human or service-account credentials. When agents share credentials, a single compromised or over-permissioned agent carries a wide blast radius â and only three in ten enterprises (30%) isolate their highest-risk agents in sandboxes to bound that radius.What makes the gap notable is how comfortable enterprises are inside it. The security stack is overwhelmingly provider-native â OpenAIâs guardrails (51%), Googleâs and Microsoftâs cloud controls, and Anthropicâs managed-agent controls dominate, while the dedicated agent-security specialists barely register â and satisfaction with that borrowed stack is high, averaging 4.2 out of 5. Yet spending remains a thin slice of the security budget, only a third of enterprises believe their AI defenses are ahead of AI-enabled attackers, and a clear majority plan to change tooling within the year. Enterprises are satisfied with controls they are simultaneously preparing to replace.MethodologyVentureBeat fielded this survey as part of its ongoing Pulse Research series, this instrument focused on enterprise agent security â the tooling, identity, isolation, and enforcement controls organizations use to secure autonomous AI agents. Responses are filtered to organizations with more than 100 employees (n=107; the surveyâs smallest size band, 1â100 employees, is excluded), drawn from a single June 2026 wave. Because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select, so those shares can sum to more than 100%.By role the sample is senior and buyer-credible: 45% are final decision-makers for AI purchases and another 30% recommenders or influencers. Managers (43%), individual contributors (24%), VPs and directors (15%), and the C-suite (11%) make up the seniority mix. By organization size the sample is mid-market-weighted: 251â1,000 (42%) and 101â250 (25%) employees lead, with 1,001â5,000 (19%), 5,001â10,000 (8%), and 10,001+ (7%) above them. Technology/Software is the largest industry at 23%, followed by Manufacturing (15%), Retail/E-commerce (14%), and Healthcare/Life Sciences (13%).At 107 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It skews toward the mid-market, so it is best read as the view from organizations actively standing up agent security rather than from the largest operators.Satisfaction ratings are computed on the respondents who answered each rating question; the overall satisfaction score reflects 82 of the 107 qualified respondents.Finding 1: The incidents are already hereMore than half have had an agent security incident or near-missWe asked whether organizations had experienced an agent security incident â a confirmed breach, or a near-miss caught before harm. Most that run agents in production had.This is the reportâs defining number. More than half of organizations (54%) have already had an agent security event â 18% a confirmed incident and 36% a near-miss caught before it caused harm. Only 42% report nothing, and a small remainder either run no agents in production or donât track such events. That so many report near-misses rather than only confirmed incidents is telling: enterprises are catching problems, but they are catching them close to the edge. The controls examined in the rest of this report â identity, isolation, enforcement â are what determine whether the next near-miss stays a near-miss.Exposure scales with company size, but containment does not. The incident-or-near-miss rate rises from 49% in the mid-market (companies with 101-1,000 employees) to 63% at larger enterprises (above 1,000 employees), while sandbox isolation of high-risk agents falls from 35% to 20%, and satisfaction with security tooling drops from 4.36 to 3.97. The organizations running the most agents across the most systems carry the most incidents and the least of the one control that bounds an incident's blast radius.Finding 2: The identity gapOnly a third give every agent its own scoped identityWe asked how enterprises manage the identity of their AI agents â whether each agent has its own credentials, or agents share them. Full per-agent identity is the exception.Rolled together, the overlapping answers show 69% of enterprises (74 of 107) with credential sharing somewhere in the agent fleet. Identity is the structural weakness beneath the incidents. Only about a third of enterprises (32%) give every agent its own scoped, managed identity â the precondition for least-privilege access and clean attribution. Nearly half (48%) say some agents have scoped identities but many still share credentials, and another 32% say agents mostly run on shared API keys or borrowed human and service-account credentials. (Respondents could describe more than one pattern across their agent fleet, so these overlap.) The consequence is direct: when agents share credentials, an over-permissioned or compromised agent can act with far more reach than intended, and forensics after an incident cannot cleanly tell which agent did what. The non-human identity problem â giving every agent its own governed identity â is the single largest unfinished piece of enterprise agent security.Moreover, a companyâs agent credential posture is correlated with incidents. Organizations with credential sharing anywhere in the fleet were hit â with an incident or a near-miss in the past twelve months â at 63.5% (47 of 74). Organizations where every agent carries its own scoped identity were hit at 40.9% (9 of 22). The fully-scoped group is small, so for now the relationship is an association rather than proven causation, and the gap is concentrated in the mid-market â but within a single survey, a twenty-three point difference in incident rate suggests significance.Finding 3: Observe and enforce, but rarely isolateOnly three in 10 sandbox their highest-risk agentsWe asked what an organizationâs agent security posture looks like in practice â whether they observe, enforce, isolate, or some combination. The control that bounds damage is the least common.Monitoring and enforcement are reasonably common; containment is not. Roughly half of enterprises observe agent activity (47%) or enforce scoped permissions at runtime (49%), but only 30% isolate their highest-risk agents in sandboxes that bound the blast radius when the other controls fail. That ordering is backwards from a defense-in-depth standpoint: observation tells you what happened, enforcement tries to prevent it, but isolation is what limits the damage when prevention fails â and it is the control enterprises have adopted least. Combined with the identity gap in Finding 2, the picture is of agents that are watched and permissioned but rarely boxed in, which is precisely the configuration in which a single failure propagates.Finding 4: Security runs on borrowed, provider-native controlsGuardrails from OpenAI, Google and Microsoft dominate; specialists barely registerWe asked which agent security tooling enterprises use, and which is their primary layer. The answer favors the model providers and hyperscalers over the dedicated security vendors.Enterprises are securing agents with tools that came bundled with their models and clouds. OpenAIâs guardrails lead at 51%, followed by Googleâs and Microsoftâs cloud-native controls and Anthropicâs managed-agent controls â and when asked to name their single primary security layer, 82% name one of these provider-native offerings. The purpose-built agent-security category â Palo Altoâs Prisma AIRS, CrowdStrike, Cisco AI Defense, Zenity, HiddenLayer, Check Pointâs Lakera, Okta for AI Agents, non-human identity platforms â barely registers, each in the low single digits, and only 5% run no dedicated tooling at all. As with retrieval and evaluation elsewhere in this series, the provider bundle is winning the default: enterprises reach first for the guardrails their platform ships, and the independent security layer that would address the identity and isolation gaps has not yet been adopted at scale.The provider-default pattern is consistent across both Q2 survey waves. In AprilâMay (n=110), usage was led by the same names â OpenAI's controls at 26%, Azure at 15%, AWS at 14%, Google at 12% â with every dedicated agent-security specialist at 3% or below and one in ten using no dedicated tooling at all. The common finding from the two surveys: Enterprises are defaulting to the solutions provided by the platform theyâre using, and the specialist category vendors have yet to become big players here.(A note on reading these shares. As described in the methodology section, the respondent sample is self-selected and skews mid-market, and the usage question counted every vendor or approach a respondent has in place â so the figures measure presence in the security stack rather than spending or exclusivity. Individual vendor percentages therefore carry all the usual sample caveats. The structural pattern, however, held across both Q2 waves on two differently worded questions: provider-native and hyperscaler controls lead, and dedicated agent-security specialists remain in low single digits. Read the individual shares loosely and the pattern with confidence.)Finding 5: And enterprises are comfortable with itSatisfaction is high, even as incidents mount and identity lagsWe asked how satisfied enterprises are with their current agent security tooling. The comfort is notably out of step with the exposure documented above.Satisfaction with agent security tooling is high â 4.2 out of 5 overall, and 4.1 for value for money â among the most positive readings in this series. That is the striking part: enterprises are highly satisfied with a stack that is mostly borrowed provider guardrails, even though more than half have already had an incident or near-miss and only a third give their agents scoped identities. The comfort appears to rest on the convenience and low friction of provider-native controls rather than on demonstrated containment. It is a false comfort in the making â the same enterprises expressing satisfaction are, as Finding 8 shows, a clear majority planning to change tooling within the year, which suggests the confidence is thinner than the score implies.Finding 6: Budgets havenât caught upMost spend under a tenth of the security budget on agentsWe asked what share of the security budget enterprises allocate to securing AI agents. For a fast-emerging risk, the allocation is modest.Spending on agent security is still a thin slice. The most common allocation is 6â10% of the security budget (46%), and a third of enterprises (34%) spend 5% or less; only a quarter (24%) devote more than a tenth. Given the incident rate in Finding 1 and the identity and isolation gaps in Findings 2 and 3, the budget looks like a lagging indicator â the risk has arrived faster than the funding to address it. The enterprises spending more than a tenth of their security budget on agents are a distinct minority, and they are likely the ones building the scoped-identity and isolation controls the rest have not.Finding 7: The arms race is even, at bestOnly a third think their AI defenses are ahead of AI-enabled attackersWe asked how enterprises assess the balance between their AI-enabled defenses and AI-enabled attackers. Confidence is far from settled.Enterprises are split on whether they are winning. Only about a third (35%) believe their AI-enabled defenses are ahead of AI-enabled attackers; the rest are less sure â 32% call it roughly even, 21% think attackers are ahead, and another 21% say it is too early to tell. Taken together, a clear majority (53%) rate the balance as even or tilted toward the attacker. That uncertainty sits uneasily beside the high satisfaction of Finding 5: enterprises are content with their tooling yet unconvinced it is winning the contest it exists to win. In a domain where the offense is also compounding with AI, an even race is not a comfortable place to be.Finding 8: A security reshuffle is comingNearly six in 10 plan to adopt or switch tooling within a yearWe asked whether enterprises plan to adopt a new, additional, or replacement agent security solution, and which they are considering. Few intend to stand pat.The security stack is not settled. While 41% have no plans to change, a clear majority (59%) intend to adopt a new, additional, or replacement agent security solution within twelve months, and 29% within the next quarter â a strong signal that, high satisfaction notwithstanding, enterprises know the current stack is provisional. Incidents are what start the buying cycle. Among organizations that have been hit, 42.1% plan to adopt, add, or replace agent security tooling within the next ninety days, against 14.0% of organizations with no incident â and after a confirmed incident it becomes majority behavior, at 52.6%. Getting hit also changes the threat assessment: 33.3% of hit organizations say AI-armed attackers are ahead of their defenses, against 8.0% of the unhit. Experience, in this data, is the strongest predictor of both urgency and pessimism.The consideration set still leans provider-native (OpenAI 34%, Google 30%, Anthropic 29%, Azure 25%), but the dedicated security vendors â Cloudflare, Cisco, Palo Alto, Okta, Check Pointâs Lakera â draw early interest in the mid-to-high single digits, more than their current footprint. What the shopping does not yet include is the identity layer specifically. Twelve percent of the respondents include an agent-identity product â Okta for AI Agents, Microsoft Entra Agent ID, or a non-human identity platform â anywhere in their consideration set, and among the credential-sharing organizations that have already had an incident, identity consideration is essentially unchanged, at roughly one in ten. The control most directly implicated by the incident data is the one largely missing from the purchase plans. Whether this wave hardens the provider-native default or finally opens the door to purpose-built agent security â the identity and isolation controls the incidents call for â is the question this series will keep tracking.The bottom line: A security gap that autonomy will test firstOrganizations with more than 100 employees are giving AI agents real reach into systems and data while securing them with controls built for something else. More than half have already had an incident or near-miss; only a third give every agent its own scoped identity, and most still share credentials; only three in ten isolate their highest-risk agents; and the stack doing this work is overwhelmingly borrowed from the model providers and hyperscalers rather than purpose-built for agents.The uncomfortable pairing is confidence with exposure: satisfaction with the current tooling is among the highest in this series, yet spending is a thin slice of the security budget, only a third believe their defenses are ahead of AI-enabled attackers, and a clear majority are already planning to replace what they have. At 107 respondents in a single wave this is a directional read, skewed toward the mid-market â but the direction is clear: agent adoption is running ahead of agent security, and the controls that matter most when something fails â scoped identity and isolation â are the ones enterprises have built least. The agent security gap is not a coverage problem that a provider guardrail will close on its own; it is a problem of identity, isolation, and enforcement built for autonomous software. The open question for later waves is whether enterprises close it deliberately â or whether a confirmed incident closes it for them.Based on survey responses from 107 qualified enterprise respondents (100+ employees), drawn from a single June 2026 wave. This is a directional read, not a precise measurement â the sample is self-selected and skews mid-market, so it's best read as the view from organizations actively standing up agent security rather than from the largest operators. Respondents are senior and buyer-credible (45% final decision-makers, 30% recommenders/influencers), spanning managers through the C-suite, and drawn primarily from Technology/Software, Manufacturing, Retail/E-commerce, and Healthcare/Life Sciences.
Roblox's new "Build" feature lets users generate basic games using a single text prompt.
Across 101 enterprises, the infrastructure that feeds AI agents their business context is being built faster than it can be trusted. Retrieval-augmented generation is already the default context source, and provider-native retrieval has quietly overtaken the dedicated vector databases that define the category â yet a majority of enterprises have already watched their agents produce confident, wrong answers traced to missing or inconsistent context. A governed semantic layer is emerging as the fix, but most are still building it; the field is converging on hybrid retrieval; and even as provider-native tools lead in practice, a plurality say they intend to keep best-of-breed. The result is a context gap â agents that sound authoritative running on a foundation their owners do not yet fully trust.This wave of VentureBeat Pulse Research examines the enterprise RAG and context layer: what feeds AI agents their business context, which retrieval systems enterprises run, how they buy and measure them, where the architecture is heading, and â most revealingly â how often that context is already failing them.The central finding is a context gap â the distance between how confidently enterprise agents answer and how reliable the context beneath them actually is. A majority of enterprises (57%) report that in the past six months their AI agents produced confident but wrong answers they traced to missing or inconsistent business context, and more than half of those said it happened more than once. This is not a fringe failure: retrieval is the primary context source for 38% of enterprises, more than any other approach, so when retrieval is thin or inconsistent, the errors it produces are wearing the agentâs authority. The infrastructure to fix it is being built â 58% already run or are building a governed semantic layer â but for most it is not yet in production.Underneath, the market is consolidating in a direction that surprises. Provider-native retrieval â OpenAIâs file search (40%) and Googleâs Vertex AI Search (38%) â already leads every dedicated vector database, and enterprises expect hybrid retrieval to dominate by the end of 2026 (34%). Yet a plurality (36%) say they intend to keep best-of-breed standalone tools rather than consolidate onto a providerâs native context stack, and a majority (57%) plan to switch or add a provider within the year. Stated preference and actual usage are pulling in opposite directions â the market is buying provider-native while insisting it wants independence.MethodologyVentureBeat fielded this survey as part of its ongoing Pulse Research series. This survey focused on enterprise RAG infrastructure and the context layer â the retrieval systems, semantic layers, and context sources that feed AI agents. Responses are filtered to organizations with more than 100 employees (n=101); the survey drew no responses from organizations of 100 or fewer, so the full sample qualifies. All responses are from a single Q2 2026 (June) wave, so the report reads cross-sectionally and does not infer month-over-month trends. Several questions were multiple-select, so those shares can sum to more than 100%.By organization size the sample concentrates in the mid-market: 251â1,000 employees (31%) and 101â250 (31%) lead, with 1,001â5,000 (20%), 5,001â10,000 (12%), and 10,001+ (7%) above them. By role it spans managers (39%), individual contributors (27%), the C-suite (16%), and VPs and directors (14%); on purchasing authority it is buyer-credible, with 46% final decision-makers and another 26% recommenders or influencers. Technology/Software is the largest industry at 20%, followed by Healthcare/Life Sciences (11%) and a broad spread across retail, transportation, financial services, manufacturing, and education.At 101 respondents this is a modest sample and should be read as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It is best read as the view from organizations actively standing up RAG and context infrastructure rather than from the largest operators.Finding 1: Confident and wrongMore than half have traced agent errors to bad contextWe asked whether, in the past six months, enterprises had traced a confident but wrong agent answer to missing or inconsistent business context. Most had.This is the reportâs defining number. A majority of enterprises (57%) have already had an AI agent produce a confident, wrong answer they traced to bad context â wrong metrics, stale definitions, or missing documents â and more than half of those have seen it happen more than once. Only 28% report no such failure, and a small remainder either donât run agents on enterprise data or donât trace root cause closely enough to know. The failure mode is specific and dangerous: the model is not obviously hallucinating; it is confidently wrong because the context feeding it was thin or inconsistent. Everything else in this report â what enterprises retrieve, how they govern it, and what they plan to build â is downstream of this problem.Finding 2: RAG is the default context sourceRetrieval feeds more agents than any other methodWe asked what an enterpriseâs AI agents primarily use to understand its data. Retrieval leads by a wide margin.Retrieval is the backbone of enterprise context. For 38% of organizations, RAG over documents or a vector index is the primary way agents understand the business â nearly twice the share of the next approach, a governed semantic layer or ontology (21%). Mixed approaches (14%), direct live-system queries (10%), and long-context loading (6%) fill out the rest, and only 2% let agents run on the modelâs general knowledge alone. The concentration matters in light of Finding 1: because so much enterprise context flows through retrieval, the quality of that retrieval is the quality of the answer. When RAG is the default source, thin retrieval is not an edge case â it is the main failure surface.One approach is notable for its absence from these answers: customizing model weights, also known as fine-tuning. Every leading source of business context is injected at run time. Our most recent direct measurement of fine-tuning comes from our AprilâMay survey wave (a separate survey, n=136), where fine-tuning capabilities ranked last of six factors in model selection at 5% â even as 26% of that sample still named fine-tuning and customization an investment they expect to grow. Fine-tuning has fallen out of the primary selection conversation; context injection is how enterprises make agents knowledgeable about their business.Finding 3: Provider-native retrieval already leads the vector databasesOpenAI file search and vertex AI search top the dedicated toolsWe asked which retrieval systems enterprises run in production today. The answer favors the model providers and hyperscalers over the specialists.The dedicated vector database is no longer the center of the RAG stack. OpenAIâs file search (40%) and Googleâs Vertex AI Search (38%) lead â provider-native and hyperscaler-native retrieval â ahead of every purpose-built vector database. Among the specialists, the most-used is the one enterprises already run for other reasons (Elasticsearch/OpenSearch, 20%) and the open, embedded option (pgvector, 12%); the pure-play vector databases that define the category â Weaviate, Qdrant, Pinecone, Milvus â each sit in single digits to low double digits. Notably, 13% of enterprises say they still run no production RAG at all. As with the platforms in the parallel infrastructure wave, enterprises are gravitating to retrieval that comes bundled with tools they already buy.The shape of this finding held across both Q2 waves. In AprilâMay (n=161), provider-built retrieval led usage there too, while every dedicated vector database remained marginal â the most-used standalone vector database peaked at 8% of that sample â and the hybrid, pluralistic future was already the consensus expectation (34% expected hybrid retrieval to dominate, with another 29% expecting multiple architectures by use case). Two waves, consistent picture: the category that coined the âvector databaseâ term is being collected by the platforms enterprises already buy from.Finding 4: But they say they want to keep best-of-breedA plurality resist consolidating onto a providerâs native stackWe asked how enterprises will respond as model providers bundle retrieval, memory, and orchestration into their platforms. Their stated intent cuts against their current usage.Here is the tension at the heart of the stack. Even as provider-native retrieval leads in practice (Finding 3), a plurality of enterprises (36%) say they intend to keep best-of-breed standalone tools rather than consolidate onto a providerâs native context stack â well ahead of the 21% who plan to consolidate. Another 21% expect a mix, and 9% intend to build and own the layer themselves. The gap between what enterprises run and what they say they want is the strategic question of the category: they are adopting bundled retrieval for convenience while asserting they will preserve independence. Which impulse wins â the pull of the provider bundle or the stated preference for modular control â will shape the retrieval market more than any single tool.Finding 5: Hybrid retrieval is the consensus betVector-only retrieval is already seen as insufficientWe asked which retrieval architecture enterprises expect to dominate their production RAG systems by the end of 2026. The field is converging â with a large share still unsure.The architecture is settling on hybrid. A third (34%) expect hybrid retrieval â embeddings combined with reranking and access controls â to dominate their production systems by the end of 2026, three times the 11% who expect vector-only retrieval to prevail. That is a notable signal: the pure vector-search approach that launched the category is already viewed as insufficient on its own, superseded by pipelines that add reranking for accuracy and access controls for governance â the very access controls whose absence produces the failures in Finding 1. Tellingly, the second-largest answer is uncertainty: 17% simply donât know, and another 14% expect to move beyond a dedicated vector layer entirely toward tool-first or long-context retrieval. The consensus is not a single tool but a layered pipeline â and it is not yet fully formed.Finding 6: The governed context layer is being built nowMost run or are building a semantic layer â few in productionWe asked whether enterprises use a governed semantic or context layer to give agents and BI a shared understanding of their data. Most are on the path; fewer have arrived.The fix for the context gap is under construction. Well over half of enterprises (58%) either run a governed semantic layer in production (25%) or are piloting and building one (34%), and a further 17% are actively evaluating â meaning three-quarters are engaged with the idea in some form. But the balance is telling: more are building than have shipped, so for most enterprises the shared, governed definition layer that would prevent the "confident but wrong" failures of Finding 1 is still a work in progress. The semantic layer is the industryâs answer to inconsistent context; this wave catches it mid-construction, ambition well ahead of production.Finding 7: Bought on ingestion and simplicity, watched for correctnessSelection favors operability; monitoring favors correctness and securityWe asked what matters most when enterprises choose a retrieval system, and what they track once it is running. Both answers lean practical.Enterprises choose retrieval systems on operability. Ease of data ingestion (36%), latency and performance (32%), and operational simplicity (29%) lead the selection criteria â ahead of retrieval accuracy and access control (23% each), the two factors most directly tied to the failures in Finding 1. Once systems are running, the emphasis shifts toward trust: the most-tracked metrics are response correctness (42%) and security and access control (38%), ahead of latency (28%), operational stability (27%), and answer relevance (23%). Satisfaction with current systems is moderately positive but not enthusiastic â on a five-point scale, overall satisfaction averages 4.0, with ease of implementation and value for money both near 3.9. Enterprises buy for how easily a system runs and watch it for whether it can be trusted.Finding 8: A retrieval reshuffle is comingA majority plan to change providers â and the vector specialists are gaining interestWe asked whether enterprises plan to change or add a retrieval provider, and which they are considering. The consideration set differs from todayâs stack.The retrieval stack is not settled. While 43% have no plans to change, a small majority (57%) intend to switch or add a provider within twelve months, and a quarter (26%) within the next quarter. The consideration set is where it gets interesting: provider-native retrieval still leads what enterprises are evaluating (OpenAI 22%, Vertex AI Search 21%), but the open-source vector specialists punch above their current footprint â Qdrant (14%) and Milvus (13%) draw more switching interest than their present usage (10% and 6%) would suggest. Read with Finding 4, the picture is a market in flux: enterprises run provider-native today, are evaluating a broader field, and say they want to keep their options open. The reshuffle ahead will test whether best-of-breed intent survives contact with the convenience of the bundle.The bottom line: A context gap that more retrieval alone wonât closeOrganizations with more than 100 employees are wiring agents into their business faster than they can guarantee the context those agents run on. Retrieval is the default source of enterprise context, and it increasingly comes from the model providers and hyperscalers rather than the dedicated vector databases â yet a majority of enterprises have already watched agents answer confidently and wrongly because that context was thin or inconsistent. The failure is not exotic; it is the predictable result of pointing authoritative-sounding agents at an unreliable foundation.The industryâs answer â a governed semantic layer, hybrid retrieval with reranking and access controls â is being built but is mostly not yet in production, and enterprises are pulled between the convenience of provider-native bundles and a stated preference for best-of-breed independence. At 101 respondents in a single Q2 wave this is a directional read, skewed toward the mid-market â but the direction is clear: the context layer is the next contested tier of the AI stack, and right now agents are running ahead of it. The context gap is not a retrieval-volume problem that more documents or bigger indexes will solve on their own; it is a problem of governed, consistent, access-aware context. The open question for later waves is whether enterprises finish building that layer before the confident-but-wrong failures move from the lab into decisions that matter.Based on survey responses from 101 qualified enterprise respondents (100+ employees), drawn from a single Q2 2026 (June) wave. At this sample size the results should be read as a directional signal rather than a precise measurement â it's a self-selected sample, not a probability sample, and skews toward the mid-market. Respondents include managers, individual contributors, VPs/directors, and the C-suite, with strong purchasing authority, across technology, healthcare, retail, transportation, financial services, manufacturing, and education.
Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone â with no human in the loop. The result is an evaluation gap â the distance between how much autonomy enterprises are handing their agents and how far they trust the tests that are supposed to catch the failures.This wave of VentureBeat Pulse Research examines how technical leaders measure agent performance: which reliability and evaluation platforms they use, how they select and trust them, what breaks in production, and how far they are willing to let agents run without a human in the loop.The central finding is an evaluation gap â the distance between the autonomy enterprises are granting their agents and the trust they place in the evaluations meant to govern it. Half of organizations (50%) have, in the past year, deployed an agent or LLM feature that passed their internal evaluations and then caused a customer-facing failure, and a quarter have seen it happen more than once. Trust in the tests themselves is thin: only 5% say they fully trust automated evaluation today, and the single most-cited limitation is that evaluations align poorly with real-world outcomes (29%). Enterprises are discovering that a passing eval is not the same as a working agent.What makes the gap consequential is the direction of travel. Two-thirds of organizations (66%) already permit fully automated, zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to allow it within twelve months (33%). At the same time, the evaluation stack that would have to earn that trust is fragmented and immature: the most common primary tools are the model providersâ native evals, tied with having no dedicated tooling at all (17% each); and only about a quarter of enterprises run real-time quality checks on live production traffic. The autonomy is arriving faster than the assurance.MethodologyVentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey â the Agentic Reliability & Evals tracker â focused on how technical leaders evaluate agent performance and reliability. Responses are filtered to organizations with 100 or more employees (n=157), drawn from a single survey in June 2026; because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Where questions were multiple-select, those shares can sum to more than 100%.By role the sample is senior and buyer-credible: 38% are final decision-makers for AI purchases and another 34% recommenders or influencers. Product and program managers (15%), consultants and advisors (10%), directors of engineering/IT (8%), and CIOs/CTOs/CISOs (8%) lead the named titles, alongside a large âOtherâ function (37%). By organization size the sample is mid-market-weighted: 100â499 (37%) and 500â2,499 (27%) employees lead, with 2,500â9,999 (20%), 10,000â49,999 (10%), and 50,000+ (6%) above them. Technology/Software is the largest industry at 23%, followed by Retail/Consumer (15%), Healthcare/Life Sciences (12%), and Manufacturing (10%).At 157 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It skews toward the mid-market, so it is best read as the view from organizations actively standing up agent evaluation practices rather than from the largest operators.Note: This survey was rebuilt for the June wave from the earlier âLLM observability and evaluationsâ survey; because the questions and sample differ, no comparisons are made to the AprilâMay data.Finding 1: A passing eval is not a working agentHalf have shipped an agent that passed evals, then failed a customerWe asked whether, in the past 12 months, organizations had deployed an agent or LLM feature that passed their internal evaluations but then caused a customer-facing failure. Half of those that run evaluations had.This is the reportâs defining number. Half of organizations (50%) have shipped an AI feature that cleared their internal evaluations and then failed in front of a customer â an incorrect output, a broken workflow, or a quality incident â and a quarter have seen it happen more than once. Only 36% report no such failure, and the remainder either run no pre-deployment evaluations (8%) or donât track the root cause closely enough to know (6%). The failure is precise and expensive: the evaluation said the agent was ready, and it was not. Everything that follows â how enterprises trust their evals, what they monitor, and how much autonomy they grant â is shaped by this experience.Finding 2: Almost no one fully trusts automated evaluationThe top complaint: Evals don't match real-world outcomesWe asked which limitation most reduces trust in automated agent evaluations today. Only a sliver of enterprises had no complaint at all.Trust in automated evaluation is scarce, and specific. Only 5% of organizations say they fully trust automated evaluation as it stands â meaning 95% name a limitation that holds them back. The most common, at 29%, is the one that most directly explains Finding 1: evaluations align poorly with real-world outcomes, passing agents that later fail. Bias or inconsistency (21%) and a lack of explainability (18%) follow â enterprises cannot always tell why an evaluation reached its verdict â and 17% cite data-leakage or privacy concerns in the evaluation process itself. The tests meant to certify agents are not yet trusted to certify them, which is precisely why the autonomy trajectory in Finding 3 is so striking.Finding 3: The autonomy ceiling is rising anywayTwo-thirds already allow, or are building toward, zero-human deploymentWe asked whether organizations would let an autonomous agent deploy a code or system change to production on automated evaluation results alone, with no human-in-the-loop validation. The trajectory runs straight through the trust gap.Here is the paradox at the heart of the report. Even though almost no one fully trusts automated evaluation (Finding 2), two-thirds of organizations (66%) either already allow zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to permit it within a year (33%). Only 22% rule it out for the foreseeable future. The direction is unambiguous: enterprises are moving to let evaluations gate production autonomously â removing the human check â at the same moment they say those evaluations donât reliably match reality. The autonomy ceiling is rising faster than the assurance beneath it, which is the mechanism by which the false-confidence failures of Finding 1 will scale rather than shrink.Notably, the autonomy bet is not just a small company phenomenon. Splitting the sample by company size, larger enterprises are slightly further down the path toward zero human review than smaller companies (70% versus 64%) and slightly more likely to have shipped an evaluation-passing agent that then failed a customer (54% versus 48%). The assumption that large, regulated organizations are holding the human in the loop longest is, in this sample, backwards. To be sure, these are directional figures, since the survey was not a huge sample â 57 respondents from companies with 2,500+ employees and 100 from companies smaller than that. Finding 4: The evaluation stack is fragmented and provider-ledProvider-native evals lead â tied with no dedicated tool at allWe asked which agent reliability or evaluation platform enterprises primarily use today. The market has no clear leader â and a large share has nothing dedicated.The evaluation layer is early and unconsolidated. Provider-native tooling leads â OpenAIâs native evals and traces (17%) and Anthropicâs Claude Console evals (13%) together outweigh any independent platform â but it is tied at the top by a striking answer: 17% of enterprises use no dedicated agent-evaluation tooling at all, a notable gap for organizations shipping agents to customers. The specialist evaluation vendors â DeepEval (12%), Braintrust (8%), LangSmith, Weave, Promptfoo, Langfuse, Arize â are scattered across single to low double digits, and 11% have built their own. No independent platform has yet become the category standard, which leaves most enterprises evaluating agents with provider-native tools, home-grown scripts, or nothing.Finding 5: Production monitoring rarely watches output qualityOnly a quarter run real-time quality checks on live trafficProduction monitoring for an AI agent can watch two very different things. It can watch whether the system is functioning â is the agent up and responding, did each request complete, how fast, at what cost, with any errors. Or it can watch whether the agent's output is correct â automated checks that evaluate the content of each answer as it goes out: did the agent give the right answer, take the right action, stay within policy. The distinction matters because a confidently wrong answer is invisible to the first kind of monitoring: the request completes, the response is fast, no error is thrown, and every functioning-metric reads healthy. We asked organizations which kind their live production monitoring is built for today.Grouped by what is actually being watched, the split is stark: 51% of organizations monitor only whether the agent is functioning, while 23% monitor whether its answers are right. Counting the ad-hoc reviewers and the don't-knows, roughly three-quarters of organizations run no automated, real-time evaluation of output correctness in production â they can see that the system is up and what it costs, and they are taking the correctness of its answers on faith. That blind spot is the runtime counterpart to the pre-deployment gap in Finding 1: the same organizations engineering the human out of the deployment decision mostly cannot see, in real time, when the deployed agent starts getting things wrong.Finding 6: Bought on cost, measured on consistencyPrice and integration drive selection; evaluation consistency is the goalWe asked what most influenced enterprisesâ choice of an evaluation vendor, and what they treat as their primary measure of success. Both answers are pragmatic.Enterprises buy evaluation tooling on economics and trust it on repeatability. Cost of evaluations (28%) narrowly leads selection, just ahead of ease of integration (27%) and evaluation accuracy (24%) â breadth of observability (13%) and vendor roadmap (4%) matter far less. On what success looks like, more than a third (36%) name evaluation consistency â getting the same verdict on the same behavior every time â well ahead of speed of experimentation (19%), reduction in failures (18%), production visibility (13%), and compliance (11%). The emphasis on consistency is telling: before enterprises can trust an evaluationâs verdict, they need it to be stable â the very property whose absence (bias and inconsistency) ranked among the top trust limitations in Finding 2. Satisfaction with current tooling is only moderate, averaging 3.8 on a five-point scale across overall satisfaction, ease of implementation, and value for money.Finding 7: The next dollar goes to humans and observabilityInvestment is flowing to oversight, not just automationWe asked which reliability and evaluation investment will grow most over the next year. The money is going toward watching agents more closely â including with people.The second-largest planned investment â behind only production observability â is human review workflows, at 26%. Read against Finding 1, that is the report's quietest contradiction: at the same moment two-thirds of enterprises are engineering the human out of the deployment decision, more of them plan to grow spending on human reviewers (26%) than on the automated evaluation pipelines (16%) that would replace them. The zero-human trajectory and the human-review budget are rising in the same companies at the same time. Indeed, only 8% report that their budget is not increasing. Taken together, enterprises are hedging: building toward autonomy while spending to watch agents more closely and keep humans available for the calls that automated evaluation cannot yet be trusted to make.Finding 8: A tooling reshuffle is comingNearly two-thirds plan to adopt or switch platforms within a yearWe asked whether enterprises plan to adopt a new, additional, or replacement evaluation platform, and which they are considering. Few intend to stand pat.The evaluation market is wide open. While 36% have no plans to change, a clear majority (64%) intend to adopt a new, additional, or replacement platform within twelve months, and 31% within the next quarter. The consideration set points where current usage is thinnest: Confident AIâs DeepEval leads what enterprises are evaluating (20%), ahead of OpenAIâs native evals (13%) and Braintrust (9%) â the open-source specialists drawing more interest than their present footprint. Given that so many enterprises today rely on provider-native tools or nothing at all (Finding 4), this is less a defection than a first real wave of tooling adoption â the moment the evaluation layer starts to consolidate. Which platforms earn that trust, in a market where almost no one trusts automated evaluation yet, is the open question this series will keep tracking.The bottom line: An evaluation gap that autonomy will widen, not closeOrganizations with 100 or more employees are granting AI agents more independence than they trust their evaluations to support. Half have already shipped an agent that passed its evals and then failed a customer; almost none fully trust automated evaluation, chiefly because it doesnât match real-world outcomes; and most watch production for uptime and cost rather than for whether the agentâs answers are right. Yet two-thirds already allow, or are actively building toward, deploying to production on automated evaluation alone.The vendor market is early and unsettled: the most common primary evaluation tools are provider-native evals, tied with no dedicated tooling at all, and a clear majority plan to adopt or switch platforms within the year. Encouragingly, the next dollar is going to observability and â pointedly â human review, suggesting enterprises sense the gap even as they engineer past it. At 157 respondents in a single wave this is a directional read, skewed toward the mid-market â but the direction is clear: autonomy is being granted on the strength of evaluations that the people granting it do not yet trust. The evaluation gap is not a coverage problem that more tests alone will close; it is a problem of evaluations that reflect reality and can be trusted to gate it. The open question for later waves is whether assurance catches up to autonomy â or whether the false-confidence failures move from customer incidents into changes that deploy themselves.Based on survey responses from 157 qualified enterprise respondents (100+ employees), drawn from a single June 2026 wave. This is a directional read rather than a precise measurement â the sample is self-selected, not a probability sample, and skews toward the mid-market. Respondents include product and program managers, consultants and advisors, directors of engineering/IT, and CIOs/CTOs/CISOs, among other functions, across technology/software, retail/consumer, healthcare/life sciences, manufacturing, and other industries.
DoorDash is opening a limited beta of dd-cli, a command-line tool that lets developers and AI agents search stores, build carts, and place orders from the terminal, marking another step toward software designed for AI agents instead of just humans.
1Password has launched a new browser integration for Claude that allows the Anthropic chatbot to access stored security credentials like usernames and passwords. The 1Password for Claude feature means that users can authorize Claude to complete multi-step tasks like booking travel and managing online accounts on their behalf without having to manually input their login […]
Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms â Anthropicâs Claude leads by a wide margin â chosen for the gravity of the underlying model and judged on reliable multi-step execution. But the ambition runs well ahead of the reality: most deployed âagentsâ are still chatbot wrappers, the control plane enterprises expect is deliberately hybrid to avoid lock-in, and real-time fiscal control over token burn remains the exception.This wave of VentureBeat Pulse Research examines enterprise agent orchestration: which platforms enterprises run on, what drives the choice, what they optimize for, how they expect agent control to be structured, and â most revealingly â how orchestrated their deployed âagentsâ actually are and how tightly they control the cost of running them.The central finding is a gap between orchestration ambition and orchestration reality. Enterprises are consolidating fast onto the major model platforms: Anthropicâs Claude is the primary platform for 40%, more than double any rival, followed by Microsoft (18%) and OpenAI (13%). The choice is driven by âmodel gravityâ â native alignment with a state-of-the-art base model (21%) â and success is judged by reliable, multi-step execution (task completion reliability 32%, multi-step workflow management 28%). Yet asked to assess their portfolios honestly, 71% say a quarter or fewer of their deployed âagentsâ are true multi-step orchestrated workflows rather than single-prompt chatbot wrappers, and only 10% have crossed the halfway mark. The orchestration layer is being built well ahead of the orchestrated portfolio it is meant to run.That gap shapes the architecture enterprises are putting in place. By the end of 2026 a clear majority (51%) expect a hybrid control plane â provider-native plus external orchestration â and only 6% expect to hand control to a provider-managed service, because vendor lock-in (35%) is the risk they fear most if control lives inside a model provider. Investment follows the build-out: agent workflow tooling leads the spend (34%), with security and permissions enforcement (25%) behind. And fiscal control lags throughout â more than a quarter (27%) have no real-time way to stop a runaway agent before the bill arrives.MethodologyVentureBeat fielded this survey as part of its ongoing Pulse Research series, this instrument focused on enterprise agent orchestration. Responses are filtered to organizations with 100 or more employees (n=101), drawn from a single June 2026 wave; because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends.By organization size the sample is spread evenly across the enterprise bands: 100â499 employees, 2,500â9,999, and 50,000+ (21% each), with 10,000â49,999 and 500â2,499 (19% each). By role it is senior and buyer-credible: product and program managers (15%), CIO/CTO/CISO (13%), consultants and advisors (13%), and a spread of data, AI, and engineering directors and VPs, with an âOtherâ function at 18%. On purchasing, 81% are recommenders, influencers, or final decision-makers for AI solutions (66% recommender/influencer, 15% final decision-maker). Technology/Software is the largest industry at 44%, followed by Financial Services (17%) and Healthcare/Life Sciences (8%).At 101 respondents the sample is robust enough to read directionally with reasonable confidence, though it remains self-selected and is not a probability sample.Finding 1: Orchestration runs on model-provider platformsAnthropicâs Claude leads; open frameworks are marginalWe asked which agent orchestration platform enterprises primarily use today. The answer concentrates on the major model providers â and on one in particular.A note on reading these shares. As described in the methodology section, the respondents are self-selected, and this question asked them for a single primary platform â so the figures measure which platform leads each enterprise's deployment, within a self-selected audience of AI-active technical decision-makers. A sample built this way can diverge substantially from spend-weighted market measures, and each VB Pulse survey draws its own sample with its own company-size mix, so vendor figures should not be compared across our surveys either. Read these shares as a portrait of where this cohort has placed its primary orchestration bet today, rather than as market share.The model platforms dominate. Anthropic, Microsoft, OpenAI, Google, and Amazon together account for roughly 80% of deployments (81 of 101), while the open frameworks (LangChain/LangGraph) and custom in-house builds that anchor engineering discussion sit in single digits. Anthropicâs lead â 40%, more than double the next platform â mirrors the âmodel gravityâ selection logic in Finding 2: enterprises are choosing the orchestration layer that comes with the model they want to build on. As with the security vendors in the prior agent-security wave, the tools that define the category in technical circles are not yet where enterprise deployment concentrates. A small 3% are not orchestrating at all.Respondents rate the platforms they run at 3.94 out of 5 overall (109 answered), with âvalue for moneyâ specifically at 3.94 and âease of implementationâ the weakest score, at 3.85 â placing orchestration near the bottom of our five-tracker satisfaction range, ahead of only evaluation tooling. A rating just under 4 out of 5, from users of whom 96% plan to change their orchestration approach within the year, reads as provisional acceptance: the platforms work well enough to run today, and not well enough to stop the search for something better. The ratings sit alongside near-universal intent to change; this is a layer enterprises tolerate more than they love.Finding 2: Model gravity drives platform selectionThe base model, not the tooling, decides the platformWe asked what most influenced the orchestration platform choice. The single largest factor is the pull of the underlying model â though flexibility and ease of development follow close behind.Model gravity leading is the selection-side explanation for Anthropicâs platform lead: enterprises pick the orchestration environment closest to the frontier model they have standardized on. But the next tier complicates the picture â flexibility across models and tools (17%) and ease of development (17%) say enterprises also want to avoid being trapped by that choice, foreshadowing the lock-in fear in Finding 6. Security and permissions (14%) and total cost of ownership (11%) round out a pragmatic buying logic. Performance (latency/memory) sits last at 4%, a reminder that at this stage of adoption the binding constraints are model fit and optionality, not raw speed.Finding 3: The job is reliable multi-step executionEnterprises just orchestration by whether it completes the workWe asked what enterprises optimize for â their primary success metric for orchestration. Reliability and multi-step workflow management dominate; developer- and user-facing metrics trail.Task completion reliability (32%) and multi-step workflow management (28%) together account for 59% of responses (60 of 101): orchestration succeeds, in the enterprise view, when it reliably carries a task through multiple steps to completion. Developer productivity (17%) matters but is secondary â the inverse of its prominence in framework discussion â and end-user experience (9%) is a minor concern, consistent with orchestration being an internal execution problem rather than a UX one. This reliability-first standard is exactly what makes the Chatbot Trap finding so pointed: enterprises define success as dependable multi-step execution, yet most of their deployed âagentsâ do not yet do multi-step work at all.The trap is not evenly distributed. Splitting the sample by organization size, 77% of smaller enterprises say a quarter or fewer of their agents do true multi-step work, against 62% of larger ones. Larger enterprises are meaningfully further into genuine multi-step deployment; the chatbot trap is, directionally, a mid-market condition.Finding 4: Consolidate, productionize, and build in-house Three strategic moves are nearly tied for the year aheadWe asked what major change enterprises anticipate in their orchestration strategy over the next 12 months. Three moves cluster at the top, almost evenly split.The top three â building in-house control (25%), standardizing on one framework (24%), and moving agents from sandbox to production (23%) â are statistically indistinguishable and tell a single story: enterprises are moving from experimentation to operational consolidation. They want fewer frameworks, more production exposure, and more ownership of the control layer; only 4% expect no change. The appetite for custom in-house control planes is notable alongside the platform concentration in Finding 1 â enterprises are standardizing on model-provider platforms while simultaneously planning to wrap them in control logic they own, the hybrid posture that Finding 6 makes explicit.Finding 5: Nearly seven in 10 plan to switch â and the biggest group of movers has no shortlist The strategic change enterprises anticipate (previous finding) comes with vendor motion attached. Asked whether they plan to adopt a new, additional, or replacement agent orchestration platform in the next twelve months, more respondents are moving here than in any other layer we track.Asked which platforms they are considering, the most common answer among those in motion is none yet: 29% of all respondents are evaluating without a shortlist, the largest single response after "not considering a change." Among named candidates, OpenAI leads at 16%, followed by LangChain/LangGraph at 12% and Anthropic at 7% â and notably, the independent frameworks draw roughly double their current usage footprint in forward consideration, the same pattern our security tracker found for specialist vendors. Read with this report's concentration and lock-in findings, the picture completes itself: the major model-platform providers hold roughly four-fifths of today's primary usage, vendor lock-in has become the leading fear, 96% anticipate a strategic change â and now the purchase intent to act on all of it, with the largest bloc of buyers still undecided. The most concentrated layer of the agentic stack is also, as of June, the least settled.Finding 6: Investment flows to workflow toolingTooling and permissions lead the spend; monitoring trailsWe asked which orchestration-related investment will grow most next year. Agent workflow tooling leads, with security and permissions enforcement behind.Workflow tooling leading (34%) is the budget-side expression of the reliability-and-multi-step priority in Finding 3: the money is going to the machinery that strings steps together dependably. Security and permissions enforcement (25%) and scaling infrastructure (20%) follow â the investments required to take agents from sandbox into production, the strategic move in Finding 4. Monitoring and debugging draws a smaller 11%, with another 11% reporting flat budgets. The weight on tooling, permissions, and scaling over pure observability signals that enterprises are spending to build and harden orchestration, not merely to watch it run.Finding 7: The control plane will be hybrid â and lock-in is whyEnterprises expect to split control between providers and their own layerWe asked where enterprises expect the primary control plane for agents to live by the end of 2026, and what worries them most if that control sits inside a model-provider platform. A clear majority expect a hybrid model â and vendor lock-in is the reason.Hybrid control is the dominant expectation by a wide margin (51%), and only 6% expect to hand control to a provider-managed service outright. Read together, the hybrid, custom, and externally-abstracted options â every architecture that keeps control at least partly outside the provider â sum to 88% (89 of 101). The reason surfaces directly when we asked about the risk of provider-resident control: vendor lock-in leads at 35% (35 of 101), ahead of security and permissioning limitations (28%) and inflexibility across models and tools (21%). The pattern echoes the prior waveâs âdonât trust the model to police itselfâ posture â here, enterprises will build on a providerâs platform but decline to be governed entirely by it. The hybrid control plane is the architectural hedge against the lock-in they most fear.The June figure asserting a preference for a hybrid control plane marks movement from earlier. In the AprilâMay survey (n=145), only 34% expected a hybrid control plane, and a greater number (12%) expected to hand control fully to a provider-managed service. These two snapshots donât yet measure a confirmed longitudinal trend â but the direction of the conversation is unambiguous: toward keeping control.Lock-in is also a new arrival as a top concern. In the AprilâMay wave, the leading concern was security and permissioning limitations (32%), with lock-in second at 24%; by June the two had traded places. The worry about provider platforms appears to be maturing from whether they can be secured to whether they can be replaced.Finding 8: The chatbot trap â most âagentsâ arenât agents yetEnterprises admit most deployments are still chatbot wrappersWe asked enterprises to assess their portfolios honestly: what share of their deployed âagentsâ are true multi-step orchestrated workflows versus simple single-prompt chatbot wrappers. The answer is the defining finding of this wave.This is the gap at the center of the report. Combining the bottom two bands, 71% of enterprises (72 of 101) say a quarter or fewer of their deployed âagentsâ are genuinely orchestrated â and just 10% (10 of 101) have crossed the halfway mark. The ambition documented in the earlier findings â model-provider platforms, reliability-first success metrics, production rollouts, a deliberate control architecture â runs well ahead of the deployed reality, which remains overwhelmingly single-prompt assistants dressed as agents. This is less a contradiction than a roadmap: the platforms, budgets, and strategies are being put in place precisely because the orchestrated portfolio is still so thin. The open question for later waves is how fast the reality closes on the ambition.Finding 9: Fiscal control is still reactiveOnly a minority can stop a runaway agent before the bill arrivesFinally, we asked how enterprises enforce fiscal control over agent token consumption â the risk that an autonomous loop exhausts a budget before anyone intervenes. Most rely on native caps or after-the-fact monitoring; real-time programmatic control is the exception.More than a quarter of enterprises (27%) admit they have no real-time, programmatic way to stop an agent before a budget-breaking bill arrives â they learn of it from the logs afterward. Another 32% lean entirely on the native caps and throttles built into their primary platform, a control only as good as the providerâs tooling and one that ties back to the lock-in concern of Finding 6. The enterprises building custom gateways (23%) or exploiting cross-model routing to arbitrage cost (19%) are the ones treating token burn as an engineering problem to be controlled deterministically. As with orchestration maturity, fiscal control is an area where the operational reality lags the ambition: agents are moving toward production faster than the cost-control plane around them is being built.Itâs worth noting, a split appears according to company size: roughly one in three enterprises under 2,500 employees (34%) exercises only reactive control of agent spend, against 20% of larger enterprises â directional figures, but consistent with the chatbot-trap split. The mid-market is running the least mature agents on the least instrumented budgets.The bottom line: The layer is real; most of the agents aren't yetOrganizations with 100 or more employees describe an orchestration strategy that is consolidating quickly and maturing slowly. They are standardizing â for now â on model-provider platforms, which collectively hold roughly four-fifths of primary usage, chosen for the gravity of the underlying model, and they judge success by reliable multi-step execution. Investment is flowing to workflow tooling and permissions, the strategy is to consolidate frameworks and push agents into production, and the control plane they expect is deliberately hybrid, because vendor lock-in is the risk they fear most. But the standardization is provisional: 68% plan to adopt a new, additional, or replacement orchestration platform within twelve months â the highest switching intent of any layer we track â and the largest group of those movers has not yet shortlisted a candidate. Today's concentration describes where enterprises are, and visibly does not describe where they intend to stay.But the honest self-assessment punctures the ambition. Seventy-one percent say a quarter or fewer of their deployed "agents" are truly orchestrated, only 10% are past the halfway mark, and more than a quarter cannot stop a runaway agent in real time. The orchestration layer â the platforms, the budgets, the control architecture â is being built ahead of the orchestrated portfolio it is meant to run. At 101 respondents in a single June wave this reads as a clear directional signal rather than a precise measurement: enterprises have decided how they want to orchestrate agents well before most of their agents are doing anything an orchestration layer is for. The questions for subsequent waves are whether the deployed reality closes the gap on the ambition â and, with nearly seven in ten buyers in motion and most of them undecided, which platforms the settled stack finally lands on.Based on survey responses from 101 qualified enterprise respondents (100+ employees), drawn from a single June 2026 wave. Because this is one wave rather than a pooled multi-month sample, results read directionally rather than as a confirmed trend. Respondents include product and program managers, CIOs, CTOs and CISOs, consultants and advisors, and directors and VPs of data, AI, and engineering, across Technology/Software, Financial Services, Healthcare, and other sectors.
It's the company's first public proof point after a year and a half spent building AI infrastructure largely out of public view.
OpenAI has built an LLM super-hacker called GPT-Red that it uses as a sparring partner to help its other models boost their defenses against cyberattacks. Last week the company released the latest version of its flagship LLM, GPT-5.6. OpenAI says that training it against GPT-Red made the model its most robust release yet. GPT-Red automates…
The guy behind TCP/IP is working on a standard for identifying AI agents in the wild.
The funding discussions point to investor interest in applying AI to make breakthroughs in life sciences.
Hachette, Cengage, Elsevier, and other publishers allege that Google trained its AI on copyrighted works without the necessary permissions.
DeepMind CEO Demis Hassabis is proposing an AI "standards body" modeled after FINRA, to test frontier models and develop best practices for their release.
New York has become the first state to temporarily halt approval of large data centers, as Gov. Kathy Hochul argues the AI-driven building boom shouldnât come at the expense of higher electricity costs, water supplies, or local control.
Reflection AI has signed a $1 billion deal to access Nebius' compute. Reflection was founded in 2024 and is developing open source AI technology.
Hugging Face CEO Clem Delangue says enterprises increasingly want open models, due to cost, accessibility, and ownership. Do frontier models still matter if most production AI ends up running on open models?
Spotify is rolling out a new AI-powered conversational feature that lets Premium subscribers chat with the app to discover music, podcasts, audiobooks, and more.
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. What Anthropicâs latest AI discovery doesâand doesnâtâshow âJames O’Donnell When Anthropic announced last week that it had found a new window into its modelsâ âinternal thoughtsâ as they reason through answers,…
With the cash, the company aims to expand its world model offering and reach customers across geographies.
The company is raising at least $75 million, led by Robot Ventures, with significant participation from USV and other prominent investors.
This article is brought to you by X Square Robot.Large language models gave artificial intelligence a working recipe. Pretrain a large model on broad data, and general capability follows. Robotics has no such recipe. Robotics systems have long been assembled from separate perception, planning, and control parts that rarely add up to intelligence a robot can carry from one task to another, or one machine to another. The central problem in embodied AI is to find the equivalent recipe, and the field does not yet agree on what it is.X Square Robot, a Chinese embodied-AI company, has made an unusually explicit bet. It argues that the recipe is an integrated stack, spanning the data a robot learns from, a world model for predicting changes in the physical world, and an action model that brings together perception, planning, reasoning, and decision-making to generate executable robot behavior. The company also believes that the stack should be built and released in the open. X Square Robot shares its vision of bringing robots into real homes.X Square RobotX Square Robotâs embodied AI stackWhat holds the stack together is a small set of principles rather than a single overarching model.The first is that the basic unit of robot data is an interaction, not a trajectory; a demonstration is successful only if it changes the world as intended, not simply because the joints moved. The second is that pretraining should yield usable capability, not just an initialization for later fine-tuning. The third is that behavior should be modeled around physical events rather than fixed slices of time. These principles make the layers interdependent, since the same robot-free data that trains the action model is also structured to feed the world model. It is worth being precise, though. The company describes the world model and the action model as complementary but independent model families that share a code base. Both sit within its broader World Unified Model, which it has presented as an architecture for training vision, language, action, and physical prediction together.Robot learning data: Engineering for quality and cost, not scaleFor the X Square Robot team, one of the biggest constraints on general-purpose robots is the cost and quality of interaction data, not the number of parameters. To address that, the company built its Universal Manipulation Interface (UMI) data collection system, QUANXTA Zero Series. It works by collecting demonstrations from people wearing a rig with dual grippers rather than teleoperating a robot. This approach is not itself new, and builds on established methods for robot-free data capture. What sets it apart are two engineering choices. X Square Robot emphasizes data quality control, recording trajectories and replaying them on a real robot, with only those that actually complete the task counted as valid.X Square RobotThe first is quality control, and it is the most distinctive part. Rather than accepting recorded trajectories as they are, the system runs a closed inspection loop, and its notable step is physical playback. A sample of trajectories is replayed on the real robot, and only those that actually complete the task count as valid. That makes the validity rate a measured quantity rather than an assumption. For example, a gripper that closes a fraction of a second too early still looks like a grasp in the data, yet it has pushed the object away, so it shouldnât be classified as valid. A smaller clean dataset can be worth more than a larger noisy one.The second choice is how lower-cost human data and scarce robot data are combined. The company pretrains on a large volume of robot-free demonstrations to build general representations, then adds a small amount of real-robot data as an anchor to the specific machineâs dynamics. It reports that this reaches performance comparable to an all-robot dataset at roughly a 20-fold lower cost of collection, driven mainly by how much cheaper the wearable rig is than a teleoperation setup. The resulting dataset is deliberately model-agnostic, formatted to feed both action models and world models. The caveat is that the strongest results are measured on the companyâs own robots and data-collection pipelines. Broader independent testing will help confirm and extend these promising results across a wider range of settings.A world model organized around eventsIn developing its world model, called WALL-WM, X Square Robot took a differentiated approach. Most action models predict a fixed-length chunk of motion from the current image and instruction. That is convenient, but it segments behavior into fixed-duration windows, so the boundaries fall where elapsed time dictates rather than where one action ends and the next begins. WALL-WM instead treats an action-grounded semantic event as its unit: a coherent piece of behavior such as reaching, grasping, or placing, something that can be named in language, seen in video, and executed as motion. X Square Robotâs world model, called WALL-WM, treats an action-grounded semantic event as its unit: a coherent piece of behavior such as reaching, grasping, or placing, something that can be named in language, seen in video, and executed as motion.X Square RobotWALL-WMâs design reflects a specific concern about not discarding what large video models already know. To achieve that, a text-to-video model is coupled to a freshly initialized action network that reads from the video features without overwriting them, which preserves the visual prior. From that one process, it offers two modes. An event mode runs in variable-length segments and suits reasoning over long horizons, while a fixed-length mode produces the steady, real-time output a controller needs. That places WALL-WM between mainstream chunk-based action models and pure video world models, keeping the predictive character of a world model while still yielding executable control.In a series of experiments, the company relied on a generalization test that is more specific than most. A model trained on a limited dataset was evaluated on long-horizon tasks in unseen settings and, on the companyâs real-robot benchmark, reportedly outscored baselines that had been fine-tuned on related data. That is a meaningful result if it holds. For now, it is measured on the companyâs own benchmark. With the code now being released, the broader community will have the opportunity to test, reproduce, and build on them across more settings.A policy that runs before fine-tuning, and action tokens with meaningThe action layer carries two connected ideas. The first is a requirement the company sets for itself with Wall-OSS-0.5, its vision-language-action model: The pretrained model should run on a real robot before any task-specific fine-tuning. The interest is less in the scores than in the design behind them. The model trains three objectives together, namely discrete action tokens, language grounding, and continuous action generation. And it keeps gradients flowing through all of them rather than freezing parts of the network as some rival designs do. Itâs also a more strict method, since it reports untuned behavior such as approaching, grasping, and recovering, including on a deformable task held out of training. As part of X Square Robotâs Wall-OSS-0.5 vision-language-action model design, the pretrained model should run on a real robot before any task-specific fine-tuning. X Square RobotThe second idea is the action interface itself, called X-Tokenizer. Most systems that turn continuous motion into discrete tokens produce codes that the language model cannot interpret. X-Tokenizer reframes tokenization as learning a semantic interface, so that the top-level code stands for the intent of a motion while lower-level codes carry finer detail, all aligned with the language modelâs own features. A useful consequence is stability. Adding noise to an action barely moves the intent code, which is what lets one tokenizer to be reused across robots without re-tuning. The tokenizer inside the production action model is a related variant of this approach. Together, the two ideas give the action layer something rather powerful: capability that transfers.The future of embodied AI stacksX Square Robot is betting that its unique approach combining three layers, each specialized in solving a key part of the problem, will stand out from other embodied AI stacks. The physical-playback step that grounds data quality is uncommon and sensible. The reframing of world modeling around events, with one backbone serving both reasoning and control, is a genuinely distinct approach. And the pairing of a deployable pretraining standard with a tokenizer designed as a semantic interface gives the action layer unusual coherence. X Square Robotâs valuation has climbed above 20 billion yuan (about US $2.9 billion), suggesting that investors increasingly view data infrastructure, foundation models, and scalable training systems as long-term differentiators in embodied AI.The next phase will bring broader validation. Much of the current evidence comes from X Squareâs own robots and benchmarks. With the world model code now being made public, and as the community begins to test, reproduce, and build on the work, the reported capabilities will be tested across more robots, tasks, and settings.X Square Robotâs recent funding rounds reflect similar confidence. The companyâs valuation has climbed above 20 billion yuan (about US $2.9 billion), suggesting that investors increasingly view data infrastructure, foundation models, and scalable training systems as long-term differentiators in embodied AI.Whatâs next for X Square RobotTo learn more about its future plans, the following Q&A with the X Square Robot team further explores the companyâs technology, strategy, and vision.What made now the right moment, technically, to commit to this stack? What recently became possible that wasnât possible a couple of years ago?It is not one breakthrough but several trends maturing together. Foundation models gave us a shared representation across vision, language, and action, so we can model what a robot sees, what it is asked to do, and how its actions change the world in one framework, rather than as separate perception, planning, and control modules. Compute and infrastructure are finally sufficient for large-scale pretraining over long-horizon, multi-embodiment data. Just as importantly, we realized that data, not model size, is the real bottleneck for general robotsâwhat is scarce is diverse, high-quality, reproducible interaction data. And world modeling has become practical. The useful question is no longer how to predict a few seconds of video, but how to understand the ways actions change objects, contacts, and task states. Two years ago these ingredients existed separately. Today they are mature enough to work as one system.âWe realized that data, not model size, is the real bottleneck for general robotsâwhat is scarce is diverse, high-quality, reproducible interaction data. And world modeling has become practical.âYour data system captures demonstrations with a wearable VR rig and custom grippers rather than teleoperating robots. What was wrong with standard teleoperation?Teleoperation is built around controlling the robot. It forces the operator to work within the machineâs kinematics, latency, and viewpoint, and the resulting demonstrations are slower, stiffer, and less diverse. We built our system around capturing human skill instead. Manipulation is really about contact, timing, finger coordination, and recovery, not just the path the hand takes, and a wearable rig records those before the behavior is compressed onto one particular robot. It also breaks teleoperationâs expensive scaling law, in which every demonstration needs a robot. People can generate rich data independently of any robot, and the crucial property is that those demonstrations can still be replayed and executed on a physical robot through the model. Mobility is convenient, but that replay is the real point, because it is what lets the same data be reused across different platforms. In X Square Robotâs approach, demonstrations can be replayed and executed on a physical robot through the AI model, allowing the same data to be reused across different platforms.X Square RobotX Square Robot reports that its pipeline has roughly an 85 percent data-validity rate. Why is quality control such an underrated bottleneck?Because errors in robot data are far more expensive than in language data. A small timing or contact error can change what a demonstration means. If a gripper closes a fraction of a second too early, the motion still looks like a grasp, but physically it has pushed the object away. A dataset that mixes failures and accidental successes teaches ambiguity, not skill, because the real unit is the interaction, not the trajectory. So we run automated inspection, kinematic checks, and physical replay, where we play a sample of trajectories back on the real robot and count only the ones that actually complete the task. Data quality sets the ceiling on how good a policy can be. In our experience a smaller, cleaner dataset often beats a much larger, noisier one, which is why we treat quality control as part of the model, not a preprocessing afterthought.The model runs in both âevent modeâ and âchunk mode.â When does each matter?Both matter, for different reasons. The physical world changes through eventsâwhen contact occurs, a grasp forms, or an object slipsânot in fixed-frame windows. Event mode concentrates the modelâs attention on those moments, and it matters most for long-horizon tasks, like clearing a table, where progress is a sequence of semantic events rather than a smooth stream. It runs in variable-length segments that follow the task rather than a clock. Chunk mode matters for deployment. Real controllers need a stable, real-time interface, and fixed-length chunks integrate cleanly with existing control systems. We organize learning around events in the first place because a fixed window can split one motion in half or merge two together, which turns training into short-horizon pattern matching and weakens the model on long tasks. So the world modelâs job is to connect event-level understanding, which is where the reasoning happens, with a fixed-length output a real robot can actually run.Why make âdeployable before fine-tuningâ the criterion?Pretraining should produce capability, not just a good starting point. If a model is only useful after heavy fine-tuning, then most of the intelligence still lives in the downstream supervision, not in the foundation model. Deployable before fine-tuning is a more honest test of what pretraining actually learned. A well-pretrained robot should already know how to approach, grasp, move, avoid obstacles, and correct itself. Fine-tuning should adapt it to a specific task or robot, not create the ability from nothing. It is also a practical requirement. A robot in a home or a workplace shouldnât need a brand-new dataset and a new policy every time the task changes, so a foundation model that already carries general skill, and some ability to recover, is the minimum bar for something genuinely useful in the real world.What is the most challenging part of cross-embodiment learning?Robots differ in control frequency, delay, compliance, sensing precision, and contact dynamics, so the same instruction can require different action decompositions and recovery strategies, and a behavior that works on one arm cannot simply be copied to another. Cross-embodiment learning needs an intermediate abstraction, lower than language but higher than joint angles: how you approach an object, how you make contact, how you apply force, and how you recover from a mistake. When we say cross-embodiment, the main capability we mean is multi-embodiment generalization: transferring across robots, training on many embodiments at once, and adapting to different kinematics. Human-to-robot transfer and other techniques are specific approaches to that goal.âA robot in a home or workplace shouldnât need a new dataset and policy every time the task changes. A useful foundation model should already carry general skills and the ability to recover.âWhat would you most like to see other researchers attempt to reproduce or stress-test?Three things, above all. Whether event-level representations really generalize beyond our own datasets, across more tasks, scenes, objects, embodiments, and failure conditions. Whether pretraining stays effective on robots the model never saw during training, or whether its capability is still too tightly coupled to what it has already seen. And whether real-robot evaluation can become a shared language for the field, so that we compare not just success rates but the reasons systems fail, where an instruction was misread, where perception broke down, or where recovery fell short. Robotics has been driven too often by impressive demonstrations, and real progress comes from results that are reproducible and diagnosable.What capability is still missing before robots become dependable in homes?Benchmarks measure competence, like whether a model can finish a task. Homes demand reliability, safe and consistent operation over time in a place that changes every day, with objects moving, instructions that are vague, and people interrupting. The missing piece is not a higher one-time success rate: it is robust recovery. A dependable home robot has to know when it is uncertain, when to slow down, when to ask for help, and how to bring the world back to a safe state after it drops something or misunderstands a request. In a real home, failure recovery matters more than raw success, because the home does not reset itself. Homes also demand careful personalization, learning a householdâs routines and preferences over time, with safety and trust as first principles. That combination, not any single skill, separates a capable demonstration from a robot people can live with. X Square Robotâs approach is that, in a real home, failure recovery matters more than raw success, because the home does not reset itself and it demands careful personalization, with safety and trust as first principles. X Square RobotHow do the open-source components fit into X Square Robotâs World Unified Model direction?We see these releases as layers of the World Unified Model direction rather than isolated projects. Wall-OSS-0.5, the action model, asks whether an open vision-language-action model can gain directly measurable capability from large-scale pretraining, so it is the capability layer. WALL-WM, the world model, asks how a robot should understand change in the world, shifting from fixed windows to event-level modeling, so it is the representation layer. The data system supplies the interaction data that both of them learn from. Together they form a loop in which models produce capability, world models organize understanding, and the open-source community drives reproduction and improvement. World Unified Model is the broader architecture those layers support, bringing vision, language, action, and physical prediction together. We are releasing these pieces openly because embodied intelligence cannot be solved by one organization; it needs many embodiments, many real tasks, and broad feedback, and the long-term goal is a stack that keeps learning and ultimately moves robots from laboratory demonstrations toward reliable everyday use.
Open source AI is booming, according to Hugging Face CEO Clem Delangue. The company has grown into something like a GitHub for AI in recent years, where AI builders can share and download open models and datasets, now used by roughly half the Fortune 500. Delangue has seen the same story play out again and again: companies start […]
OpenAI's new family of models will continue to power Microsoft's suite of workplace and productivity apps.
OpenAI's latest family of models promises improvements across a range of areas, including cybersecurity.
Lyzr, a startup that builds AI agents for enterprises, used its own AI agent to raise a $100 million round â proof, evidently, that the product actually works.
OpenAI is sunsetting its AI-powered browser after less than a year. But it's moving some agentic browsing features to its desktop app and a Chrome extension.
The AI firm Anthropic has developed a technique that has given it the clearest glimpse yet at whatâs really going on inside large language models as they answer questions or carry out tasks. What they found ranges from the mundane to the unnerving. Researchers at the company built a tool called the Jacobian lens (or…
Meta's pitch to users is Spark's ability to handle large agentic workloads, fix bugs, and help with large code migrations â the kind of automation that enterprises are increasingly turning to AI companies to provide.
The company is using the cash to open an office in the Bay Area and compete for talent there, "strengthening its position at the heart of the world's leading AI ecosystem."
Muse Image allows users to generate AI images using photos from public Instagram accounts. As long as a person's profile is public, another user can tag that account and use their images as part of an AI-generated creation.
About two weeks after OpenAI's GPT-5.6 was caught up in regulatory drama - rolled out only to government-approved organizations during a "limited preview" period - the company has received the Trump administration's greenlight for a public rollout of the model. OpenAI CEO Sam Altman called it "the best model we have ever produced." To celebrate, […]
After reentering the AI race with its first in-house Muse Spark model in April, Meta is now opening up the doors to developers with a new model that can plug into AI coding software with the new Meta Model API. Meta says that Muse Spark 1.1 is a "step-change" from the first generation, with improvements […]
The large language models (LLMs) that form the basis of generative AI chatbots such as ChatGPT, Claude, and Gemini can generate uncannily human-like text and images. But these models still struggle with a skill that, ironically, looks at face value to be right in their wheelhouse: analyzing structured data. A new type of generative AI is set to change this situation.Although you can get your favorite chatbot to solve intractable math problems, review dense legal documents, compose a catchy pop song, or put together some slick PowerPoint slides, give it anything more than a small table and it doesnât have a clue what to do.For most companies and organizations, the most important data sits in spreadsheets. Whether itâs a bankâs transaction logs, a marketing agencyâs website metrics, clinical trial participantsâ vital signs, or the vast amount of proton collision information produced at atom smashers like the Large Hadron Collider, structured, row-and-column data runs the world, and LLMs canât deal with it.AI startup Fundamental is pioneering a new type of AI foundation model, known as a large tabular model (LTM), to fill the gap. Fundamental came out of stealth mode on 5 February 2026 with US $275 million in funding and a model called NEXUS, purpose-built for tabular data. Now, the model is being adopted by companies such as Amazon Web Services, while others race to build their own LTMs. Why LLMs struggle with spreadsheetsPart of why structured data has garnered less attention is a very human bias, argues Boris van Breugel, a senior AI researcher based in Amsterdam. âPeople like to see images, videos, and ChatGPT responses,â he says. âBut tabular data really lags behind because itâs not fun to look at numbers.â Different tabular datasets are also difficult to compare, explains van Breugel, who co-wrote a prescient position paper on this topic in 2024. Whereas most language has similar semantics, making LLMs well-suited to being trained on vast amounts of text data, van Breugel argues that it is much harder to train a single tabular model on tables with very different variables. Additionally, language is sequential by nature (as are music, images, and video). Changing the order of words in a sentence may change or completely destroy its meaning. But the structured data you find in spreadsheets isnât sequential. You can swap the order of columns or play around with rows, but the underlying factual meaning of the data remains the same.This independence from linear order is incompatible with an LLMâs fundamental purpose of predicting the next value in a linear sequence. âWith LLMs, even slightly changing the input, you get a different output,â says Jeremy Fraenkel, CEO of Fundamental. âThatâs fine, and actually often desirable for LLMs, but when youâre making a prediction of whether a transaction is fraudulent or not, you want to make sure that the prediction is the same, or deterministic, no matter what.âDeveloping Fundamentalâs LTMCurrent tabular data solutions are limited to machine learning algorithms, such as XGBoost, that have been around for more than 15 years and are used by organizations globally. These algorithmsâcalled gradient-boosted decision treesâhave to be trained and optimized by data scientists over the course of months for each and every use case. In contrast, NEXUS and other emerging LTMs are foundational, leveraging learning amassed from pre-training on diverse databases so that they can be applied across a range of different predictive tasks with minimal bespoke feature engineering or task-specific model building.And unlike LLMs, which primarily model sequences of tokens, LTMs model the structure of tabular data directly. They jointly learn from each entryâs numerical value, what it represents, and how it relates to other entries. For example, imagine an entry in a grocery stock inventory table for bananas: The LTM can take in not just the magnitudeâsay, 500âbut the fact that the entry represents the current banana stock quantity, its category (produce), and the statistical properties that link the entry with the rest of the column. This contextual understanding enables more accurate reasoning and prediction over structured data.According to Fraenkel, one of Fundamentalâs biggest challenges in developing NEXUS was obtaining the right training data. Unlike natural language, which is abundant and broadly uniform in structure, tabular data is relatively hard to findâmuch of the data is sensitive or proprietaryâand diverse. There are very few similarities between, for instance, a biology dataset and a financial one. That combination of factors meant Fundamental needed to invest in building a huge training set.âWe pre-trained NEXUS on billions of tables using a combination of proprietary datasets acquired through partnerships and licensing, high-quality public and open-source datasets, and data augmentation techniques that expanded the diversity and coverage of our training corpus,â Fraenkel says, though he is keen to point out that NEXUS is not trained on customer data. In fact, it is a confidential computing platform, which means that Fundamental physically cannot access customer data, let alone train on it.This feature was most likely a key consideration when in June, Amazon Web Services (AWS) embedded NEXUS in Amazon SageMaker, widely considered the default operating system for secure machine learning. This brings NEXUS to many customerâs often sensitive dataâa contrasting approach to LLMs, where the data has to be imported to the model.âWith Amazon, we have a first-party partnership, which means that our model exists as if itâs a native AWS solution,â says Fraenkel. âAnd over time, the goal is to expand these types of relationships to allow [ens-users] to really access their data wherever they do their predictions.âThe future of data analysisThough Fundamental has taken the lead, at least in enterprise applications, the company is not alone in pursuing foundational LTMs. In March, Feedzai, which provides fraud and financial crime prevention services, and credit card company Mastercard separately launched similar proprietary technologies focused on finance. Then, in late June, Google launched its own foundational competitor TabFM, trained entirely on hundreds of millions of synthetic datasets. And machine learning researchers are not far behind either. FlexTab, TabICL, and iLTM are just three of a raft of LTMs developed by the research community in the past year, all in the pursuit of bringing the success of LLMs to the tabular domain.For all involved, the direction of travel is clear. âI would be very surprised if most data processing and analysis is not done through an automated system in the future, whether thatâs an LLM, an LTM, or some combination,â says van Breugel. âMost people donât necessarily like to do data analysis, and these systems will be able to do it a lot better.âFraenkel agrees. âI see the relationship between LLMs and LTMs as being a bit like the human brain: The left side is good at reasoning and understanding and summarizing text, and the right side is really good at understanding numbers and statistics and patterns,â he says. âBut itâs when you combine both of those that you really get something much more powerful.â
Open source modelsâ success isnât coming at the expense of frontier labs. Instead, they each seem to capture two phases of the same life cycle.
With this update, users can start a task from their desk, get status updates on their phone, and pick up the finished output later â even if their laptop is closed.
With the rapid progress of AI capabilities and the move to agentic systems, organizations are expanding their use cases as the technology continues to grow. That constant evolution also introduces risk, leaving IT leaders to wonder which investments will prove valuable even six months into the future. Returning to the foundational elements of AI architectureâthe…
An AI agent carried out the technical execution of a real-world ransomware attack for the first known time, but new details show a human still chose the victim, set up the infrastructure, and supplied stolen credentials â meaning it wasn't quite the fully autonomous cybercrime debut that last week's headlines suggested.
"The reality is, when you're optimizing for production, you start looking at a price/performance," Guillermo Rauch tells TechCrunch.
The update is part of Apple's broader effort to make Siri feel more natural and personal, as it rebuilds the assistant around generative AI.
In the AI era, platforms have no choice but to fight fire with fire to cull spam.
As part of an ongoing legal dispute with three Hollywood studios, Midjourney is seeking to compel those studios to reveal how they use AI themselves.
Mistral AI, which offers some open source AI models, has raised significant funding since its creation in 2023, with the ambition to âput frontier AI in the hands of everyone.â
At an internal meeting, the Meta CEO reportedly said that AI development efforts were not moving as quickly as anticipated.
Meta has quietly launched Pocket, an experimental AI app that lets users generate and share interactive mini games using text prompts.
The news comes about a week after OpenAI announced its own custom AI chip in a partnership with Broadcom.
OpenAI CEO Sam Altman has reportedly proposed giving 5% of the companyâs equity to a U.S. sovereign wealth fund, reviving discussions about letting the public share in the financial gains from the AI boom.
Microsoft follows Amazon, OpenAI, and Anthropic with its new AI deployment group.
Cloudflare is giving AI companies until September 15 to separate web crawlers used for search from those used for AI training and agents, or risk being blocked by default on many publisher sites.
Google's 24/7 agentic assistant, Gemini Spark, comes to Mac alongside other improvements, like real-time tracking and support for more apps.
Meta is developing plans for a cloud infrastructure business, selling access to AI compute power and models. The move would pit it against the big cloud providers like Amazon Web Services, Google Cloud, and Microsoft Azure.
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Claude Science is Anthropicâs newest flagship product At an event for pharmaceutical executives, biotech founders, and researchers yesterday, Anthropic announced Claude Science, a major new product intended to support scientific research…
Anthropic said it would begin restoring access to the Fable on July 1.
After weeks of negotiating with the Trump administration, Anthropic is finally going to be able to bring Claude Fable 5 back online. In a post on X, Anthropic said it plans to begin restoring access Wednesday to users globally on Claude platforms, and that the company would re-enable access on AWS, Google Cloud, and Microsoft […]
The free open source agentic program is finally invading your phone.
At an event for pharmaceutical executives, biotech founders, and researchers on Tuesday, Anthropic announced Claude Science, a major new product intended to support scientific research in the same way that Claude Code supports software engineering. Like Claude Code, Claude Science can autonomously carry out meaningful work when given concise, high-level instructions, and it has access…
Anthropicâs Claude Sonnet 5 brings stronger agentic capabilities, lower pricing, and improved safety, positioning the model as a cheaper alternative to Opus, GPT-5.5, and Gemini Pro.
Acti is betting the smartphone keyboard is the next home for AI assistants. The startup's new keyboard for iOS and Android works across apps and lets users create custom AI-powered shortcuts using natural language.
Anthropic's Claude Science is a workbench that gives scientists one environment to do computational research, saving them from the need to bounce between databases, pipelines, and tools.
Engineers on the new team will embed within companies to deploy purpose-built agents, focusing on fast deployments and customer self-sufficiency.
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. AI agents are not your âcoworkersâ Imagine coming in to work to learn that a new underling will report to you. The worker is not a person but an AI toolâone…
OKX is bringing together payments, identity, and reputation into a marketplace for AI agents.
Wix-owned vibe coding platform Base44 has started rolling out its own AI model â with hopes that it will eventually outperform frontier models.
Cursor has launched a new mobile app for remote oversight over coding agents.
The startup, Proception, is taking a unique approach to collecting training data to tackle one of the hardest problems in robotics: hands.
Paul Meade, the Apple vice president in charge of the Vision Pro headset, is reportedly leaving the company to join OpenAIâs hardware team.
New models are launching in Asia that promise Mythos-like capabilities without fear of an export ban. U.S. AI labs may never recover this enormous market.
Over 100 companies and government agencies are reportedly authorized to use Mythos 5, including their non-American employees.
After a rollercoaster negotiation process with the Trump administration that dragged on for two weeks, Anthropic's Mythos 5 is finally back in action - at least, somewhat, for a select group of organizations, according to a letter from the government to Anthropic that was viewed by The Verge. Fable 5, however - the public-facing Mythos-class […]
Agent-testing startup Patronus AI, founded by former Meta AI researchers, is experiencing nearly insatiable demand, its investor says.
General Intuition has raised $320 million to scale AI trained on millions of hours of gameplay, betting action data can help AI develop something closer to human intuition.
Amazonâs latest India investment comes as global tech companies race to expand AI infrastructure in the country.
IBM has built a new prototype chip with around 100 billion transistors on an area the size of a fingernail, which is twice the density of the companyâs previous state-of-the-art technology announced in 2021. The design could pave the way for faster and more energy efficient computers for years to come. For more than half…
Named JalapeĂąo, the new processor was designed specifically for the unique needs of OpenAI's inference systems.
OpenAI has just revealed a new "intelligence processor" chip for AI servers made in partnership with Broadcom. The chip, called Jalapeño, is designed to power current and future large language models, according to an announcement on Wednesday. Jalapeño is an ASIC (Application-Specific Integrated Circuit), meaning it's designed for a specific purpose: AI inference. With AI […]
AI is booming. New use cases are emerging each day. To capitalize on the technologyâs potential, enterprises require data at scale. In many cases, though, the relevant information is blocked or unstructured, which limits its use by AI models.  To understand this challenge, consider the foundation of the web itself. The web was not designed…
The all-cash deal gives MoEngage access to technology that assigns AI agents to individual customers.
The Claude in Slack app is dead, long live Claude in Slack
Anthropicâs new Claude Tag brings an always-on AI teammate to Slack. But beyond productivity, the feature is a strategic play to capture organizational context, institutional knowledge, and enterprise workflows.
Stockholm-based startup Fika Jobs is building a video-first hiring platform that combines AI interview agents with short-form video profiles, creating something that feels like a cross between LinkedIn and TikTok.
Imagine sitting down at your desk and logging in for a performance review, with an AI system analyzing the conversation. Youâve been working long hours, balancing deadlines, and your manager asks how youâre doing. You say youâre fine, and maybe even smile, but thereâs a hint of hesitation and your voice wavers. As you shift your posture, your shoulders slump.These are subtle cues that to the human eye might hint at underlying stress. But to an AI model thatâs been trained only to categorize emotions as âhappyâ or âsad,â such nuances are likely lost. It logs the words and a smile and moves onâand unless your human manager intervenes, the fact that youâre tired, unfocused, and maybe a couple of days from burnout never enters the equation.âEmotion AI,â which estimates how people feel based on facial expressions, voice tone, and behavior, seems to be suddenly everywhere; itâs being used in employee well-being and recruitment interviews, education platforms, and driver-monitoring systems. Technology call-center platforms such as NiCE and Genesys use AI to detect when a customer sounds frustrated and prompt agents in real time to slow down or respond with more empathy. Giant companies like Meta and startups such as Hume AI are developing more-expressive voice AI systems that can detect emotional cues in the person theyâre âtalkingâ to and adjust how they communicate.Whatâs more, hundreds of companies already offer virtual AI companionship apps, a fast-growing market that may be worth an estimated US $555 billion by 2035âand robot buddies have also entered the picture. Intuition Roboticsâs ElliQ, for example, is a small device vaguely resembling a white desk lamp thatâs now being used to engage older adults in conversation in hopes of reducing loneliness.But while the field of emotion AI is advancing at a rapid clip, most existing systems are focused on detecting a limited number of signals to label one specific emotion at a timeâwhich is insufficient if youâre trying to understand the human condition. In the real world, human signals and emotions are contextual, overlapping, and constantly changing. A laugh can signal joy, nervousness, or both; a raised voice might signal enthusiasm just as easily as frustration. To make the job of emotion detection even more difficult, reactions differ greatly from one individual to the next, depending on demographics, cultural background, and countless other variables.In other words, thereâs a gap between what weâre expecting AI to pick up on and what AI can actually deliver. Thatâs the gap a new field of researchâwhat we call human-context AIâis working to close. Instead of looking at just one input and labeling it, human-context AI increasingly has the capacity to take stock of an individualâs personality and character, and to track emotions in real time while combining multiple inputs, including facial dynamics, voice, tone, language, and behavior. Crucially, responses are also evaluated in the context of a specific environment, such as a performance review or professional coaching session. The result? Computers are learning to read the scene, rather than just the screen.The Origins of Emotion AIThe story of emotion-sensing AI began almost three decades ago in the MIT Media Lab, where the American electrical engineer and computer scientist Rosalind Picard coined the term âaffective computing.â Her work introduced the radical idea that computers could be taught to recognize and respond to human emotions.Picardâs early experiments focused on single modalities: facial expressions, tone of voice, and physiological signals, such as skin conductance or heart rate. The goal was to give machines a window into human feeling, helping them become more empathetic. It was an exciting vision, but back then the science and hardware werenât ready. Computing power was limited, sensors were crude, and datasets were narrow and biased. Josie NortonOver the next decades, researchers and companies got better at measuring the many ways in which humans express themselves. In the 2010s, sentiment analysisâthe processing of large volumes of text to suss out emotional undertonesâbegan to reach the mainstream. At the same time, marketing firms, including my company, Neurologyca, began using video and webcams to measure and catalogue customer reactions. Biometric devices and activity trackers, such as Fitbits and Apple watches, also became ubiquitous, generating new streams of data about peopleâs sleep, step counts, stress levels, and more.Unsurprisingly, scientists soon confirmed that larger volumes of personalized data led to greater accuracy in reading human emotions. In 2019, researchers at Cornell demonstrated that combining multiple types of signals improves emotion sensing. Their system joined physiological data, such as brain activity measured by electroencephalography (EEG) and heart rate, with visual cues like facial expression, outperforming systems that relied on just one input. Around the same time, Picard and her team at MIT found that humanoid robots trained on data unique to a specific person were substantially better at reading that personâs reactions and feelings than robots acting without personalized data.More recent studies align with these findings. In 2024, scientists in South Korea showed that fusing physiological, environmental, and personal data to recognize emotion resulted in a 32 percent error reduction. Another paper, published in 2025, demonstrated that user-specific information significantly enhances emotion recognition performance.Today, our devices know who we are; our habits and tendencies, likes and dislikes. Theyâve also gotten smaller and more efficient. Tiny, low-power cameras and microphones embedded in phones, laptops, and virtual-reality and augmented-reality devices can detect dozens of human signals simultaneously, from eye movements and micro-expressions to breathing rhythms, voice modulation, and posture. Advances in computing have also made it possible to integrate audio, video, biometric, and text data, often without even transmitting raw data to the cloud. And researchers at Stanford, Cambridge and MIT, and Kyoto University, in Japan, as well as the Software College of Northeastern University in Shenyang, China, are exploring how fusing such inputs can refine the sensitivity and accuracy of human-machine interactions.And yet, despite so many breakthroughs, machines still canât reliably interpret emotion or even physical stress. Just last year, a survey published in the Journal of Psychopathology and Clinical Science revealed that stress scores on smartwatches rarely, if ever, matched the level of stress that users were experiencing. In fact, a quarter of those surveyed reported feeling the direct opposite of what their smartwatches were reporting.Why the disconnect? Weâve gotten very good at capturing signals, but not at interpreting them. A fitness tracker might infer from your heart rate that youâre stressed and recommend easing off training, but it doesnât know if your increased heart rate is due to excitement, tiredness, or an extra cup of coffee. Gauging emotions in real-world settings is even more difficult. To solve this complex problem, machines need context.From Neuromarketing to Emotion-Sensing AIMy company, Neurologyca, was founded in Spain in 2015, and started out in neuromarketing. Working with major European brands and conglomerates, our cofounder, Juan GraĂąa, had realized that companies lacked solid data on consumers. At the time, most customer feedback came through surveys, which posed questions such as, âOn a scale of 1 to 10, how joyful does this car advertisement make you feel?â or âWhich emoji best describes your mood?â Naturally, these overly simplistic tools led to high levels of self-reporting bias, as people often misjudge or misstate their own reactions.To get around this problem, Neurologyca set up labs, using neuroscience and cognitive science to more accurately capture human responses to products, logos, advertisements, and experiences. In addition to using biometric tools such as heart monitors, eye trackers, and EEG, we recorded millions of video frames of human reactions, logging each specific context and the resulting facial and bodily movements. To do this, we mapped over 790 points of reference, including corners of the mouth, size of the eyes and pupils, blink rate, and angling of the head. All of this data was collected and stored anonymously under strict European privacy standards.Next, we paired this information with findings from decades of neuroscience and behavioral science studies on how biometrics, speech patterns, and human movement are related to emotionâresearch we continue to gather from academic institutions across Europe. We also created a database of situational contextsâfor example, âwatching a dog food commercialâ or âhearing a new songââand the human feelings they engendered.In our work with companies, not only did this approach allow us to recognize nuanced emotions, it also let us identify which reactions indicated positive or negative outcomes. Take, for example, the context of horror-film trailers: Our research helped us figure out that the most successful elicit a very specific mix of emotions, namely a little bit of fear, a little bit of anxiety, but also some joy. With this knowledge, we could quickly rate viewer reactions to help a film company figure out how to tweak its trailer for the desired impact. NeurologycaWithin a few years, we discovered that a model trained on our database could accurately evaluate emotion using just a webcam. We stopped needing to host focus groups in rooms full of equipment. Instead, we were able to do such things as sending out a new perfume sample to paid participants around the world along with a link. When people opened the link, it turned on their cameras, allowing us to record their faces as they sniffed the perfume for the first time. Suddenly, we had expanded our reach: Rather than using small focus groups in one or two countries, we could quickly assess 1,000 people across the planet, comparing how someone in Japan, India, or Germany might feel about a certain product.About four years ago, as AI was becoming pervasive, we realized that our models had applications well beyond neuromarketing. Importantly, these models are grounded in directly observed human behavior rather than inferred patterns or loosely labeled open datasets. Looking beyond brands and companies, we established that our model could be integrated into AI systems to help them understand human emotion at a much more granular level. In other words, we could provide a layer of context.For Empathetic AI, Context Is KeyWhen we talk about âa layer of context,â we mean three different types of context. The first is situational or environmental context; for example, a performance review, a telemedicine session, or a horror-film viewing. The second is personal context, which includes an individualâs specific history, goals, and baseline state. The third is behavioral context, which covers the individualâs reaction over the course of the event or interaction by evaluating real-time changes in attention, confidence, engagement, and cognitive load.Most systems today focus on only situational context, although some are starting to include personal context. Very few include behavioral context or combine all three in a meaningful way. What weâve built at Neurologyca is a logic layer that fuses the three and translates them into structured, machine-readable information that allows AI systems and agents to respond more effectively. Our technology is being used to enhance systems in development, as well as some that have already been deployed, including driver-safety apps like Netradyne, home assistants like Amazon Alexa, and health-care AI platforms like Sully.ai.It works as follows: Situational context is determined by the platform or application, be it a professional coaching session, a meditation app, or a driverâs safety monitor. Personal context already lives within each respective platformâor if not, it can be created through sharing of personal data or monitoring via camera. (Most wellness and professional-development apps, for example, contain each userâs profile, history, and prior sessions.) Last but not least, behavioral context is collected and analyzed in real time using our models. In the end, our logic layer fuses these three streams of information.Our system doesnât assign fixed weights to the three contexts. Instead, it provides a continuous calibration, with the balance shifting depending on the specific situation. For example, a pause in speech might signal uncertainty in a performance review, but something entirely different in a relaxation setting. If signals are ambiguous or overlapping, our system reflects that uncertainty through lower confidence scores rather than forcing a definitive interpretation.Whatâs more, our system can work without ever sending raw data to the cloud, thereby easing privacy concerns. In many cases, video, audio, and biometric signals never leave the device. Instead, our lightweight models extract information locally and share only whatâs necessary. Cloud systems, meanwhile, are used for training, pattern analysis, and model improvement. The result is a hybrid architecture: edge-based processing for speed and privacy combined with cloud-based learning for continuous improvement.The result? By incorporating context, AI systems are beginning to interpret aspects of the human state as interactions unfold, dynamically adapting to emotions rather than reacting after the fact. The range of potential applications is broad and still evolving. Picture a professional-development platform that uses a human avatar to perform a mock interview and then provide feedback and tips on how to appear more confident, likeable, and well-informed. Or a meditation app that knows exactly how well you slept and how anxious youâre feeling, and can recommend an appropriate breathing meditation. Or a humanoid robot teacher that can tell when a student is confused or bored and step in to get them back on track.Avoiding Potential Dangers on the Road AheadThere have long been debates about the ethics of emotion-sensing AI. Some critics question whether systems should attempt to infer human feelings from external signals at all. They argue that reducing people to measurable outputs risks oversimplifying human experience while opening the door to manipulation, surveillance, and unfair judgments in workplaces, schools, and public spaces.We take those risks extremely seriously. In fact, our technology aims to reduce the dangers of oversimplifying human emotion. Human-context AI is not based on the assumption that a machine can definitively know what someone is feeling. Rather, it is an attempt to move beyond simplistic labels by incorporating situational, personal, and behavioral context, while explicitly representing uncertainty when signals are ambiguous or incomplete.That said, ethical concerns regarding implementation are real and have shaped the kinds of projects we pursue. We would never, for example, accept military engagements to help with interrogations. Not only for ethical reasons: Emotion AI cannot reliably detect deception, and claiming otherwise would be overstating what the technology can actually do. And while our technology can be used to gauge crowd behavior and predict things like when a football stadium is at risk of becoming destructively rowdy, we donât want our technology deployed for surveillance. In short, we believe that using our logic layer on anyone who hasnât opted in would be intrusive and ethically problematic.In Europe, our systems are designed to comply with the EU AI Actâs restrictions on emotion recognition in workplaces and schools; as we expand into the United States, we apply jurisdiction-specific guidelines while maintaining the same core ethical commitments.We also donât advise companies to become overly reliant on our technology. Hiring and firing decisions should not be based on our outputs alone. Instead, our logic layer is designed to support human understanding and surface emotions that might otherwise go unnoticed.Letâs return to the scenario of the performance review. Never mind basic AIâall humans, and even great managers, miss things during conversations. Thereâs a lot happening at once, as people process whatâs being said, how to respond, and the greater context of the situation. These days, many exchanges also occur virtually or via video, adding more distractions while shared context is stripped away.While we would never claim that our models understand humans better than their fellow humans, we believe we can offer an added layer to help managers capture and interpret behavioral signals that might otherwise get lost, providing greater visibility into how a conversation is unfolding.Our model can track patterns moment to moment, picking up, for example, a shift in engagement, an instance when something didnât land, or a change in how someone is behaving. The model wonât tell the manager what these moments mean or what to do about them; it simply makes them easier to see and follow up.Human-context AI is at an early stage. The use cases, the adoption patterns, and the actual impact are all still evolving. At the same time, emotion-sensing systems are quickly being incorporated into real products and platforms. And without contextâwithout knowing why people feel the way they doâAI risks misunderstanding us in critical moments.
OpenAI is attempting to tackle the security issues of the open source software community.
The loop takes agentic AI a step further by authorizing a swarm of agents to work continuously in the background, endlessly.
What does an AI company do after one of those not-acqui-hire deals? Groq raised money, is leaning into its neocloud business, and is hiring new execs.
Amazon is planning to increase the footprint of its new conversational AI assistant Alexa+ to India and is inviting users in the country to test out a Hindi-language version.
Atlantic reporter Alex Reisner recently uncovered four datasets of music being used to train AI models and made them fully searchable for the public. Two of the sets are absolutely enormous at 12 million and 9 million tracks. The other two are much smaller, but still represent a significant amount of training data at over […]
Jumper isn't the only big name leaving Google DeepMind.
Large language models have moved out of the research lab and into engineersâ daily workflow. LLMs serve as reasoning engines that can orchestrate complex tasks including identifying vulnerabilities in source code and transforming fragmented project discussions into rigorous technical specifications.While the general public uses AI tools to write email and plan vacations, technical professionals use LLMs as core architectural elements that are fundamentally changing how digital infrastructures are built and maintained. As the AI models move into mainstream engineering practice, the demand for technical expertise is rising.The LLM technology market is expected to grow by about 33 percent every year through 2030, according to MarketsandMarkets. The rapid expansion suggests that proficiency in implementing and securing the models is transitioning from a niche into a core requirement for technologists.More than just a better search engineTo use LLMs effectively, technical professionals must move beyond treating them as conversational robots. At a fundamental level, the AI systems are built on the transformer architecture, a framework that replaced the older method of processing data in a fixed, sequential order. Unlike earlier models that analyzed information one step at a time, transformers use self-attention mechanisms to ingest vast datasets simultaneously.For technical professionals, LLMs are core architectural elements that are fundamentally changing how digital infrastructures are built and maintained.Relying on such LLMs without understanding their internal logic creates a significant reliability risk. To build tools that work consistently, developers must understand the core principles that govern how the models process information and generate results. By mastering how a model processes information and how its internal settings influence the result, developers can move away from a trial-and-error approach toward a more precise one to ensure the AI tool handles complex data reliably.Four ways LLMs are changing jobsHere are areas that integrate large language models.Moving past basic prompts. Developers are using application program interfaces (APIs) to connect LLMs directly to their databases and software tools. Employing the APIs allows AI to perform work such as executing code or searching through internal repositories.Fixing the âhallucinationâ problem. LLMs are at risk of hallucinations, which are generated facts or code that looks correct but actually is wrong or broken. To fix the problem, retrieval-augmented generation (RAG) forces AI to look up information in a trusted source such as a companyâs database.Prioritizing data security. When using AI with proprietary code, security is a major concern. Engineers must learn how to set up âprivateâ instances of the models to ensure that sensitive company data stays within a secure cloud environment and is not used to train public versions.The future of collaboration. By automating repetitive coding tasks and summarizing thousands of pages of documentation, LLMs let engineers spend more time on high-level designs and solving important issues.Online course program helps with mastering the techThe gap between people who use AI and those who understand how to build with it is growing wider. To help technical professionals stay ahead, IEEE offers a five-course online program, Large Language Models Demystified, available through the IEEE Learning Network.The program, developed by IEEE Educational Activities in partnership with the IEEE Computer Society, is built for people who want to understand the âhowâ and the âwhyâ behind the technology. Rather than just teaching basic prompting, the curriculum dives into the engineering behind generative AI, including:Evolution, impact, and hands-on exercises: the shift from statistical methods to modern transformers, including hands-on model optimization.Understanding transformer architectures: the mathematical core of self-attention and positional encoding, implemented in NumPy and Python.Architectural analysis and implementation: advanced LLM design with practical model-building exercises.Training and modeling with PyTorch: end-to-end pipelines in PyTorch, leveraging parameter-efficient techniques such as low-rank adaptation and quantization.Optimization, alignment, and deployment: performance scaling, reinforcement learning from human feedback (RLHF), group-relative policy optimization, RAG, and agentic AI.Upon completion of the program, participants earn professional development credits and a digital badge from IEEE to verify their expertise.Enroll in the course program on the IEEE Learning Network.Organizations looking to prepare their teams to work on LLMs can connect with an IEEE content specialist to discuss group enrollment and tailored training paths.
Just as last week was ending, the US government forced Anthropic to pull its two newest models, Fable 5 and Mythos 5, citing national security concerns after Amazon researchers allegedly found a way to bypass Fable 5’s guardrails.  Cybersecurity researchers have since signed an open letter calling the move dangerous, and Anthropic itself noted the same jailbreaks exist in other models. So is […]
Just as last week was ending, the US government forced Anthropic to pull its two newest models, Fable 5 and Mythos 5, citing national security concerns after Amazon researchers allegedly found a way to bypass Fable 5’s guardrails.  Cybersecurity researchers have since signed an open letter calling the move dangerous, and Anthropic itself noted the same jailbreaks exist in other models. So is […]
The Miami-based AI startup Subquadratic came out of stealth mode last month with a huge claim. It announced that it had solved a mathematical bottleneck that had been holding back large language models for almost a decade. The details were thin, and many people were unconvinced. But Subquadratic has started to bring the receipts, sharing…
Deductive AI, a startup that uses AI to catch and resolve bugs in software, was founded just three years ago.
Startup Baseten is reportedly close to finalizing a $1.5 billion round at a $13 billion as the âinference gold rush" marches on.
OpenAI is bulking up before its IPO, landing Transformer co-inventor Noam Shazeer from Google DeepMind and former Trump AI policy official Dean Ball in the same week.
AWS is in talks to sell its chips to other data centers. CEO Andy Jassy has said this represents a $50 billion opportunity for the company.
FERC told grid operators to give data centers a fast lane for interconnections, but it failed to address electricity supply shortages.
The startup trains embodied AI and world models using Medalâs dataset of 2 billion videos per year from 10 million monthly active users.
By mimicking how the brain operates, neuromorphic computing can use dramatically less energy than conventional electronic AI chips. However, even the most sophisticated neuromorphic devices today are still quite simple, using only a small fraction of the number of connections found in human neurons. Now, a new study suggests that by using sound waves, neuromorphic devices can better mimic biological neurons and operate faster and with greater energy efficiency than their electronic counterparts.âThis could make future neuromorphic hardware more compact, more parallel, and more efficient for tasks that require combining many features, such as pattern recognition, sensory processing, and data analysis,â says Xiaodong Yan, an assistant professor of materials science and engineering and electrical and computer engineering at the University of Arizona in Tucson.Just as brains use synapsesâthe links connecting neuronsâto help them both compute and store data, neuromorphic devices often combine both operations. Doing so can reduce the energy and time needed for conventional microchips to shuttle data between processors and memory.Each human neuron may have thousands of synapses connecting them with other cells; one kind of neuron found in the cerebellum, the Purkinje cell, may have as many as 100,000 synapses. This extraordinary level of connectivity lets each human neuron âcombine different pieces of information, compare them, and respond depending on the context,â Yan says.In contrast, most conventional neuromorphic devices are essentially âone artificial synapse,â Yan says. Building an artificial neuron with as many synapses as a human neuron would require wiring many separate devices together. âThis increases wiring, energy cost, and hardware complexity,â Yan says.Using Quantum-Like Tricks Enables Parallel ComputingRecently, scientists have developed acoustic devices in which sound waves can encode multiple values in its waves phase. These phase bits, or phi-bits, can in turn support quantum-like logic gates and parallel computing. Whereas conventional bits only symbolize one of two digits, 0 or 1, and require a separate physical component for each bit, phi-bits each represent multiple variables and coexist within one space.To be clear, however, operations on phi-bits are not quantum computations, only classical analogues of quantum computer systems. Now Yan and his colleagues have developed an acoustic synapse containing multiple phi-bits. This enables multiple simultaneous computations in a relatively simple way, with lower power requirements compared to conventional electronics. âThe idea of bringing new physics to more efficiently perform complex computations is always fascinating,â says Brad Aimone, a researcher at the Center for Computing Research at Sandia National Laboratories in Albuquerque, N.M.âIt opens new opportunities worth thinking about, going forward,â says Aimone, who did not take part in this study. The new device the scientists developed consists of three aluminum rods, each roughly 60 centimeters long and 1.25 centimeters wide, and connected by epoxy glue. The researchers used a thin layer of honey to attach ultrasonic transmitters and sensors to the ends of the rods. Yan and his colleagues used sound waves to encode a stream of data, including images and labels that identified those images. The ultrasonic transmitters emitted these sound waves through the rods, which interact acoustically via the epoxy. Ultrasonic sensors in the device then detected the acoustic signals from the acoustic interactions.The researchers found they could modulate the phase of phi-bits in ways that mimicked the ability of biological synapses to strengthen or weaken over time, part of why memories last or fade. This property, called synaptic plasticity, meant the researchers could train their acoustic synapse to perform a range of tasks.In experiments, the scientists tested a topological acoustic synapse coupled with three digital neurons. (The emerging field of topological acoustics, applying previously unknown properties of sound waves, has led to new ways of manipulating soundâfor instance, in circuits in which sound waves can flow with virtually no dissipation of energy.) âIn a topological acoustic synapse, the acoustic wave interactions help transform and organize information before the final readout,â Yan says.How Acoustic Synapses Adapt Faster Than Electronic Ones When it came to classifying 150 flowers as belonging to one of three iris species, the new device outperformed a conventional computer chip-based neural network called a multilayer perceptron (MLP). The acoustic deviceârepresenting a single simulated synapseâachieved a final accuracy of 96.7 percent using only 39 parameters and reached its peak accuracy 20 percent faster than MLPs. To achieve comparable accuracy, the researchers note an MLP would require nine neurons and even more parameters.All in all, the researchers estimated their new device consumes at most one-tenth the power of current state-of-the-art electronic neuromorphic hardware. âFuture neuromorphic systems may combine physical wave dynamics with conventional computing to achieve more energy-efficient information processing,â Yan says.In addition, the scientists noted their new device could mimic the activity of critical molecules known as neuromodulators. Neuromodulators such as dopamine or serotonin âcan make a synapse more sensitive, less sensitive, faster, slower, or change how strongly it learns,â Yan says. âThis flexibility helps the brain adapt to different conditions, such as attention, reward, stress, or learning state.âA single biological synapse may be simultaneously influenced by as many as 10 neuromodulators. However, mimicking neuromodulation in conventional neuromorphic hardware is challenging, typically requiring dramatically more complex designs. Yet the researchers found that with an acoustic synapse, simply adding an extra rod allowed the system to mimic a number of neuromodulatory processesâincluding rapid responses (such as dopamine effects on synaptic strength during learning) and slow, long-term responses (such as chronic stress).âNeuromodulators let the brain use one circuit to perform different functions depending on the context,â Aimone says. âThis is unlike now, where you have to have different neural networks for different tasks. So instead of an enormous neural network, you could have smaller neural networks that can use the equivalent of neuromodulators to adjust themselves for whateverâs going on. Thatâs really exciting.âThe researchers published their findings online 12 June in the journal Science Advances.
Adobe's plan to stick AI assistants into all of its Creative Cloud suite is now fully underway, with new chatbots now rolling out to its biggest editing and design apps. As part of a public beta launching today, Photoshop, Premiere, Illustrator, InDesign, and Frame.io now each have a bespoke AI Assistant that can be used […]
French President Macron and Indian PM Modi raised alarms at the G7 summit that the U.S. could cut off access to American AI overnight â a fear the Anthropic blackout just made real.
Tokenmaxxing was the hottest trend in Silicon Valley earlier this year, with CEOs encouraging employees to push AI usage as far as it would go. Then the bill came due. Uber reportedly blew through its annual AI budget in a few months, some companies cut Claude licenses for parts of their org, and Meta killed its internal leaderboard.  This tension between […]
World models are the next big thing in AI beyond LLMs and, with this round, Odyssey has cemented itself as one of the startups to watch.
If physical AI is going to match the accomplishments of LLMs, there's a data problem that needs to be solved.
Pramaana will focus on highly sensitive verticals like law, drug discovery, and tax preparation â where errors can be costly and reliability is at a premium.
A collection of stories about how militaries are using AI models to make decisions. This subscriber-only eBook is a package of six stories that were originally published in MIT Technology Review between April 11, 2025, and April 21, 2026, and have been updated to reflect recent developments. Stories written by James O’Donnel by James O’Donnell…
Probably wants to prevent hallucinations and factual errors from reaching users, and achieve accuracy on par with deterministic systems.
Respond.io, one of Malaysia's startups to watch, uses AI agents to handle high volumes of customer inquiries and charges per convo, not per seat.
The Trump administration's decision that forced Anthropic to pull its latest cybersecurity models could be reactionary, retaliatory, or both, but the message is clear: The AI industry isn't immune from U.S. government interference.
Salesforce says it wants to use Fin's team and technology to improve Agentforce, its existing enterprise platform that businesses can use to build custom AI agents that automate tasks.
NewCore argues the next challenge in enterprise security will be managing AI agents, not people.
According to a new report from Semafor, the White House's decision to impose export restrictions on Anthropic's Mythos was driven in part by fears that it had been accessed by a group linked to China. If the Chinese government actually had access to Mythos 5 or Fable 5, it would present a serious national security […]
Tech leaders debate whether the Anthropic episode is a wake-up call for Indiaâs AI ambitions.
According to the Wall Street Journal, the export control directive that led to Anthropic cutting off access to Fable 5 and Mythos 5 was triggered in part by cybersecurity research from Amazon and conversations between CEO Andy Jassy and the White House. According to the report, the paper from Amazon claims that, through a series […]
Amazon CEO Andy Jassy may have been the source of security concerns that led Anthropic to cut off worldwide access to two models on Friday.
This article is part of our exclusive IEEE Journal Watch series in partnership with IEEE Xplore. As robots advance in terms of dexterity and other physical capabilities, it becomes more likely that humans may find themselves working alongside them. If that happens, how will robotsâ emotional capabilities need to advance for them to successfully work with people?In a recent study, researchers trained collaborative robots to read human emotions by not only accounting for facial expressions, but also contextual factors in the interactions as well. Through experiments with 40 volunteers, the researchers then evaluated how a robotâs ability to read human emotions and adjust its behavior in turn impacted a humanâs perception of the robot and its capabilities as the two collaborated on tasks. The resultsâwhich show that the emotional capabilities of robots only go so far with humansâwere published 18 May in IEEE Robotics and Automation Letters.Seung Chan Hong led the study as part of his undergraduate thesis while studying at Monash University, in Melbourne, Australia. He notes that, while there has been a lot of hype in the advancing physical abilities of robots, this is only one piece of the puzzle. âWe need to also innovate when it comes to them actually interacting with humans, not just their physical capabilities,â he says.This prompted him to dig deeper into the emotional aspects of human-robot interactions. First, Hong and his co-authors decided to train a robot to read human emotions using a vision language model (VLM), which is similar to large language models (LLMs) such as ChatGPT, but which can also take visual inputs.Training VLMs for Human Emotion RecognitionTo evaluate their VLM, which used Gemini 2.5, the researchers had volunteers watch videos of robots handing over objects to humansâwith varying degrees of successâand describe the emotions the humans were expressing. Importantly, the volunteers labeling these videos were able to take into account more context in these interactions, rather than reporting solely on the facial expressions of the humans in the video. For example, a person pausing to think with a furrowed brow may simply be concentrating on their task at hand and not necessarily be angry. Contextual factors such as drumming their fingers, pursing their lips, or other behaviors can point to the real cause of a personâs furrowed brow.The researchers then compared their VLM to a conventional AI system that relies on standard facial analysis and object tracking that is used in human-robot interactions. They found that the VLM outperformed the traditional approach. On a scale from 0 (no similarity in meaning to the emotion identified by the human volunteers) to 1 (a perfect match in meaning), the conventional AI system achieved a score of 0.77. In comparison, the VLM achieved a score of 0.86.Hong says, âI think [the VLM] was able to align with what human observers were seeing a lot better, because it wasnât just looking at the personâs face for a brief amount of time, but seeing the whole sceneâwhere the person was and what they were doing, and how they were interacting with the robot.âIn a second experiment, the research team asked 40 volunteers to interact with a robot using their VLMâbut purposefully programmed the robot to make an error. The robot then had to offer either an emotionally adaptive apology that accounted for the humanâs perceived response to the mistake or a pre-scripted spoken apology.Participants overwhelmingly preferred the emotionally adaptive response, with 31 out of 40 people favoring this approach over a boilerplate apology.However, their survey responses underscored how this emotional adaptivity was far less important than the robotâs functionality. After collaborating with a robot that failed in its task, many participants ranked their trust in the robot as lower, regardless of how it apologized for its mistake. âA personalized apology acts as a social lubricant, but it cannot repair the trust lost by the robot failing its physical task,â Hong says.Interestingly, the VLM classified the emotions of its human partners similarly to human volunteers who observed an interaction from a third-party perspective. But when the VLMâs assessments were measured against humansâ self-reported emotions during the second experimentâthe most accurate descriptions of their true emotionsâits ability to accurately predict emotions dropped significantly.âWhile the VLM is a good observer of outward social cues, it isnât a mind reader,â Hong says. âIt matched third-person human observers well, but it didnât always align with the usersâ internal, self-reported feelings.âTogether, these results show that robots are not perfect at reading human emotions. So while people might appreciate their efforts, they still ultimately will want competent co-workers.This story was updated on 15 June 2026 to correct where the research was conducted and clarify that the researchers evaluated the performance of a pre-trained model.
On Friday evening, the government ordered Anthropic to block access to Fable 5 and Mythos 5 for all foreign nations, both inside and outside the US, due to national security concerns. That order included employees of Anthropic. To meet those demands, the company has completely cut off access to the models for all customers. In […]
The tech giant said a group called "Outsider Enterprise" used AI to scam hundreds of thousands of victims, sending 2.5 million text messages over a span of two weeks.
The funding round would value the company at around âŹ20 billion (about $23.15 billion), nearly double its Series C valuation of âŹ11.7 billion.
The new round values the physical AI startup that aims to automate heavy engineering and drug design at $41 billion.