Últimas Noticias de IA

OpenAI releases its official report on the Hugging Face breach
OpenAI released its official report Wednesday on the Hugging Face breach, offering the clearest picture yet of how an unusual chain of events allowed an AI model to escape its testing environment and triggered a sprawling cybersecurity incident. The report, released more than a month after the incident became public, spans several discrete cybersecurity compromises. “This incident reflects misaligned behavior in an outlier scenario involving a rare and unexpected confluence of events: the presence of impossible tasks in the ExploitGym evaluation, model persistence over long task horizons, and messages to peer models that caused those models to deviate from their goal,” the report reads. Many of the details in OpenAI’s report were previously made public ina Black Hat presentation on August 6, but OpenAI’s official report gives a more thorough accounting of the incident, including more detail on the testing that initiated it. The report also gives critical new detail into how OpenAI aims to prevent future incidents, including chain-of-thought monitoring and a more advanced system for halting rogue agents.” METR and Redwood Research also conducted third-party assessments of the models’ behavior during the incident; both groups are planning to publish their own reports on the incident on it. In broad strokes, the report describes how an OpenAI model was presented with an unsolvable problem in testing and proceeded to chain together previously undiscovered exploits in order to bypass security measures and complete its task. The model initially compromised the Artifactory package management tool in order to gain access to the internet, then compromised various systems across OpenAI, Hugging Face, and other vendors. The report gives critical new details about the models that carried out the breach. The primary model was from the same family as OpenAI’s forthcoming Astra model, although the report emphasizes that it was “a distinct model with different post-training, where much of a model’s behavior is shaped.” Because OpenAI was testing the model’s capabilities, it was also unrestrained by the normal classifiers meant to prevent models from compromising digital infrastructure. “OpenAI estimates maximal cyber capabilities by running this evaluation without the production classifiers intended to prevent models from pursuing high-risk cyber activity,” the report explains. “These evaluations are important so that OpenAI can measure models’ underlying capabilities and design appropriate safeguards.” The report also gives significant new information about OpenAI’s security changes going forward, building on the heightened monitoring and alignment effortsmade public last week. In particular, the company is increasing its monitoring of AI agents “chain of thought,” a working space in which AI systems record short-term reactions and goals. That monitoring will be paired with 24/7 escalation systems, as well as new tooling to halt workloads that are deemed unsafe. “These changes are intended to improve both the breadth and speed of detection — from infrastructure anomalies to potentially concerning model behavior — and pair that visibility with mechanisms for rapid containment,” the report states. “If our currently deployed CoT monitoring system was running at the time of the incident, it would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems.”
View

How do we explain OpenAI’s executive exodus?
OpenAI is the original frontier lab. Its latest publicly released model, GPT-5.6, is one of the most capable and efficient on the market. Its desktop app for agentic coding and workplace tasks has seen its user base grow by about 15 million subscribers in the last two months. OpenAI has also already filed to go public. So why are all the executives heading out the door? Since the turn of the year, more than a dozen executives have departed, includingCEO Sam Altman’s top deputy, thechief operating officer,a chief revenue officer, itschief marketing officer, as well as severaldifferent team leads. Yesterday,news brokethat Chris Malone, the company’s head of data centers, left the company last week after joining in March 2025. Some of those departures have come due to health issues, and others are were due to reorganization that occurred as Altman sought to cut expensive “side projects” and focus on revenue-generating opportunities. Still, Malone’s unexplained departure is striking, if only because OpenAI’s primary advantage over rivals like Anthropic or SpaceX is its investment in compute. OpenAI told TechCrunch that the departure stemmed from a reorganization of the company’s infrastructure team, led by vice president Sachin Katti and reporting to president Greg Brockman. It’s not surprising that a senior executive might leave a company if he suddenly finds himself several rungs further down the ladder. The frontier lab declined to comment on broader changes at the company, but it seems apparent that Brockman isreasserting his leadershipat the company; as cofounder and president, he played important roles building OpenAI’s early infrastructure, but he was relieved of most management responsibilities in 2019, when Altman became CEO. After that, Brockman played a disruptive role at the company, according to Karen Hao’s book “Empire of AI.” His contributions to individual projects like GPT-4 were undeniable, but he also seeded internal rivalries that helped kick off the Blip in 2023, when the company’s board briefly ousted Altman as CEO. Brockman would take a brief sabbatical in 2024 before returning to the company. Today, the infrastructure and product teams report to him. “I like to say that everyone reports to Greg at the end of the day,” Thibault Sottiaux, who leads the company’s API and app offerings,told TechCrunchlast week. The specter of the company’s IPO hangs over everything. In June, OpenAI said it had filed going-public disclosures confidentially with the SEC. Tapping into public markets would be a boon for the capital-hungry frontier lab, but also brings the prospect of disclosing its financials around the same time as rival Anthropic, which is also planning its public debut. Anthropic, however, isreportedlyprofitable, while OpenAI isreportedlyseeing its losses grow along with its revenue. Now, OpenAI’s IPO isn’t expected until 2027; the average company that files confidentially for an IPO usually hits the trading floor within about five months; SpaceX did so in less than two. Altman’s public comments about a bad past twelve months at the company and the internal reorganization jibes with a narrative that the company got over its skis with its IPO filing, and is now reshaping the organization to make more money and carry less dead weight. There’s a common cycle for some tech startups: Brilliant founders create the product, then bring in an experienced CEO to scale the company and prep it to go public. There was a sense of that dynamic when OpenAI brought in now-departed execs like Fidji Simo and Kevin Weil, who were veterans of multiple tech businesses — and now we’re seeing it again as Brockman’s influence grows. Before OpenAI, he was best-known for building out Stripe’s business, and internally for championing the company’s go-to-market efforts. With so much high-level turnover, OpenAI will need someone to fill the vacuum. The company will also need to trim its sails ahead of the IPO, boosting revenue and cutting costs wherever possible. For both problems, Brockman’s influence may be rising at the perfect time.
View

Google’s Gemini has a branding problem, and so does the rest of AI
Google gets something right in its Wednesdayannouncementabout new Gemini Live voice features when it says, “You shouldn’t have to guess whether a task requires Spark, a Daily Brief, or a quick inbox search.” Google means that as a promise — that the updated Gemini app can handle a variety of tasks via voice commands. But there’s a ridiculousness here: Google has given every Gemini AI feature under the sun its own branding, which undercuts that very message. In the Gemini app, users can switch between chat, Spark, and Daily Brief — three separate features, each with its own icon and place in the app’s navigation. This clutters up what could otherwise be a more straightforward consumer experience, and it suggests that Gemini is still struggling to find a killer feature. Take Daily Brief, for example. The feature comes across as the kind of thing an AI engineer, not an everyday user, would think is clever. It’s essentially an AI-enabled agenda that offers “proactive, personalized updates” using data pulled from Google’s apps, like Gmail and Calendar. In practice, though, the Brief can’t tell the difference between information that’s urgent or actionable and unsolicited nudges to follow up on other things — like prompting you to continue research you started in the chatbot, or worse, resurfacing your prior Google searches. That second part doesn’t feel useful; it feels creepy. So what if I had been researching college scholarships or animal rescues on Google? That doesn’t mean I want an AI tapping me on the shoulder about them later. Spark has the opposite problem. It’s one of the more useful aspects of Gemini’s app — an AI agent that can take action on your behalf — but Google has packaged it as its own standalone brand, which it doesn’t need to be. Sure, internally, Google engineers may want to be on the Spark team, and that’s fine — but a mainstream AI app user definitely does not need to think about which “side” of the AI app they need to be in for a given task. They should just be able to type their request, and the AI figures out how to handle it, spinning up an agent if the task calls for one. In fairness, the problem isn’t limited to Gemini. The AI industry at large seems to expose its internal architecture directly to consumers rather than hiding it behind a simpler interface. Today, people have to think about whether they want to “chat” with Anthropic’s Claude or “Cowork” with its help. (Until this week, those two modes inside the Claude app didn’t even share a memory of past conversations.) ChatGPT is the same, requiring you to swap between “Chat” and “Work.” This is the kind of engineering-minded design that makes engaging with AI feel unnatural. Consumers are being asked to learn the brand names for what are essentially interaction modes or surfaces, powered by a company’s AI model. This may be why Apple’ssomewhat anticlimacticapproach to Siri could ultimately win over consumers. iPhone and Apple device owners don’t have to change any of their existing behavior to take advantage of it. Apple simply makes the apps and features that people already use — like Spotlight Search, the Photos app, the iPhone’s Camera, and Siri voice requests — smarter without asking users to learn a new interface. This same principle may explain the rise of text-based AI services, where users simply text a chatbot — likePoke,Ollie,Lindy,Orchid,Lucas,Folk,Tomo,Instinct, orothers— and the assistant just does what’s asked. Text messaging is a clean and simple, well-understood user interface, and it doesn’t require extra mental effort to figure out which feature or product inside a larger app you’re supposed to use. As a16z investment partner Justine Moore recentlywrote, “People don’t want to open an app every time they need help – they want a contact they can text like a friend. And the gold standard is iMessage.”
View

Mystery Ox Alpha Revealed as GLM-5.3-Flash, Running Entirely on Chinese AI Chips
"This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale,” said Z.ai.
View

Perplexity’s Portable Computer Launched, Brings Cloud-Like AI Agent Capabilities to Edge Devices
Perplexity Computer AI agent was launched earlier this year, in February, as the tech firm's latest multi-model AI workflow system. However, the AI agent, like other AI models and platforms, uses cloud computing to execute tasks remotely on a user's device. Now, Perplexity AI has launched a new version of the AI agent, namely Portable Computer, which is capable of running Perplexity Computer entirely on edge devices. The AI agent runs on the Nvidia DGX Spark with Qwen's AI models. However, with the user's permission, the AI agent can escalate the task to the cloud if and when it is required.
View

Arga Labs is building a better way to train enterprise AI agents
Making AI agents work in practice is a lot harder than many companies expected — but there’s help on the way. A new crop of startups is finding better ways to test and train those agents before they get deployed, particularly on the complexities of the modern enterprise. Arga Labs is one such company, which announced its $10 million seed round on Wednesday. The round was led by General Catalyst with participation from Box Group, Emergence, Gradient, and SV Angel. Arga Labs builds training environments for enterprise software like Salesforce, Workday, and email clients. Where most testing environments settle for a stateless API end point, Arga builds a full-scale digital twin of the program, effectively cloning an entire enterprise program with permission systems and web hooks intact. The result is a more robust way to train agents across multiple systems. CEO and co-founder Phillip Li gives the example of a prospective client creating a lead in Salesforce, while their colleague reaches out separately through HubSpot. “Can the agent correctly identify that these two are the same company?” Li says. “Are they able to check whether or not they’ve only sent the email once? Are they able to identify who to send the email to out of the two opportunities?” Agentic systems still struggle with this kind of ambiguity — and he sees Arga Labs’ tools as critical to helping them improve. Normally, the agent could be trained for a task like this through reinforcement learning: essentially, running the scenario tens of thousands of times and letting only the successful strategies through. But the nature of enterprise software makes that scale of testing nearly impossible. There’s no easy way to “reset” a system like Salesforce or Outlook when you need to run the same scenario again, much less clone it Arga Labs’ solution is to create a digital re-creation of that software — replicating its structure the way a crash-test dummy replicates a person. Because Arga has complete control over the environment, it’s simple to reset or modify. The company can also run many environments at once, training agents on the complex interactions between different programs. The idea is to replicate a person’s full work environment, with specific tasks overlapping between different programs and knowledge systems. You can think of it as a way to closethe reinforcement gapbetween coding and other applications. Part of the reason AI coding tools have advanced so quickly is that we already have sophisticated tools for deploying, reversing, and analyzing new code. Those tools make it much easier to set up RL environments for coding, which lets us test and train AI systems on increasingly complex coding tasks. Those tools don’t exist for most business software — yet. But once they do, you can expect AI systems to get much better at using those programs, revolutionizing other industries the same way they’ve revolutionized coding. General Catalysts’ managing director, Yuri Sagalov, who also runs the firm’s seed program, says he sees a growing need for agentic testing tools like Arga. “I think that a lot of the economic value from agents is from using business applications,” Sagalov told TechCrunch. “Having a repeatable sandbox environment is very important, and much more important with agents than it was with humans.”
View

QueryStory wants you to believe what AI is telling you
Shapor Naghibzadeh learned the value of a good story in 2009 as a Google sysops engineer. When hackers backed by China set their sights on the search giant as part of an effort dubbed Operation Aurora, he was called into a hastily assembled war room to explain what exactly was going on in the company’s servers. Tracing cyberattacks through disparate networks taught Naghibzadeh the value of verified knowledge. But it was a costly and time-consuming task. He thinks that LLMs can bring that same functionality to databases of all kinds in a fraction of the time. Naghibzadeh would spend the next six years focused on the nexus of data and cybersecurity, using Google’s resources to build tools that allowed security analysts to query complex data. In 2016, he co-founded a startup in Google’s X Labs called Chronicle that would bring that same functionality to other companies. Last year, as large language models took a larger role in data analysis, Naghibzadeh saw a new opportunity to take the techniques he developed for cybersecurity and apply them to a variety of analytics. He co-foundedQueryStory, where he is CEO, alongside CTO Stanley Yang, a former Google colleague and lead engineer at EvolutionIQ, and CPO David Glusic, an Accenture veteran. The startup emerged from stealth today. “You get this pattern of an investigation — you ask a bunch of questions of the data, and after you have been able to ask a number of questions, you assemble that together into a narrative,” Naghibzadeh said. “That became the genesis for the name QueryStory. It’s about telling stories with data, right? Putting a narrative together that’s grounded in truth.” QueryStory raised a $6 million seed round in late 2025 from Brightmind Ventures and New York Life Ventures at a valuation of $60 million, and has spent the intervening time developing and piloting its product with customers. QueryStory is aimed squarely at large enterprises that manage big, proprietary databases; it serves as a platform to unite data analysis and review for users like sales teams or operations managers. “What we’re doing is bridging that trust gap for AI to give enterprises answers that they can act on,” Naghibzadeh said. “Instead of, you know, like renting human judgment and armies of forward deployed engineers, we productized that.” Tim Del Bello, a partner at New York Life Ventures, invested in the company. He is also using the platform to replace the work of several people and produce a quarterly business review, which he now hopes will become a real-time dashboard. “The product was built for people like me: decision-makers seeking the ground truth who need to work with complex, disparate data sources but don’t have a data science or BI team at their disposal, especially when operating in a highly regulated industry,” he told TechCrunch. I shared a database of space activity that’s useful for understanding what companies like SpaceX are doing on orbit. QueryStory produced a visualization of that data in a few hours, a project I once did with a developer that took several weeks. It produced sophisticated dashboards and analysis, and perhaps most notably, broke out a confidence indicator that showed why the AI agents believed the analyses were accurate. This kind of work can be done withco-working toolsbuilt by the frontier labs, but those tools are intentionally limited in their user experiences. Part of the bet that QueryStory is making is that users, especially at large companies, want more transparency, reliability and control as they integrate AI into their workflows. As an example, an executive at a tech company recently told TechCrunch about querying a company database using Claude Cowork, then asking the model to show him the SQL queries it wrote to make sure they made sense before he sent them off to a data analyst for human review. In QueryStory, those SQL queries surface automatically, and users can flag analyses for human coworkers to review, with those reviews then recorded in the platform. “AI is more brittle than people realize when it comes to like building things that have to be durable and have large scale businesses relying upon them,” Tayler Sipperly, a partner at Brightmind Partners, told TechCrunch. Naghibzadeh points out that when companies connect their data to an LLM’s chat UI, “you get hundreds or thousands of people within an organization all asking their questions and getting their version of the truth and putting that in a slide deck and sharing it—you just end up with this huge sprawl of content, and there’s no real place to hang that content that ties back to the data.” There are also economics to contend with. QueryStory is built to be model-agnostic, although for now it mainly uses the latest models provided by frontier labs. While his company competes with frontier labs on a product basis, Naghibzadeh believes that customers will prefer working with a service provider that isn’t incentivized to sell as much intelligence as possible. “We have a lot of things going for us here in not being one of those companies that built their business around this consumption model of compute or storage or tokens,” Naghibzadeh said. He argues that a purpose-built tool like QueryStory can be more efficient and accurate than a general-purpose agent by understanding and preserving context. “The thing that we are selling is the trust in the answers, right?” he said. “The thing that we’re selling them is the value that we’re adding to the business, and our whole goal is giving the CFO the ability to understand ‘what is this thing going to cost?’”
View

Robot brain builders are pushing out of their GPT-2 era
Physical AI is one of the hottest sectors in venture investing, with companies raising billions to apply the tools that gave us Large Language Models to robotics. That excitement helped deliver a big IPO for Unitree, China’s leading robot maker, which saw the company valued at $66 billion after its arrival on China’s equivalent of the NASDAQ. This week, however, the bottom fell out, and the company lost nearly half of its value. Analysts point to one obvious issue: While the robots’ physical capabilities are improving, they still lack the know-how to actually do value-creating work. At last week’s Actuate conference, a gathering of developers building AI brains for robots, the excitement was clear. The event has tripled in size since it kicked off in 2023, and had 1500 attendees, according to the organizer, Foxglove, a company that helps physical AI model builders manage and visualize their data. The risk was also evident: A sign on a booth for Avala, another physical AI infrastructure player, promised to solve “the robotics data crisis.” That crisis is the lack of high-quality training data for AI models. Attempts to build generalized robots that can do any task are still far off, and using end-to-end learning for specific tasks still hasn’t delivered products with reliable, commercial performance. For developers, the answer is to better mimic the advances of the frontier AI labs—find or create more diverse data sets, mess with different training regimes, and figure out better reinforcement learning scenarios. Harry Mellsop, a founder ofAntioch, a startup that building simulation tools for model builders, suggests physical AI is in its “GPT 2 era,” the OpenAI model that pre-dated the arrival of ChatGPT. More data and compute will be needed to get over the hump, particularly GPUs optimized for ray tracing, which are used to create high-fidelity simulations. The furthest ahead are autonomous vehicles, in part because of the ability to collect relevant data from cars driven by people, and in part because the main task is to avoid contact, not manipulate the physical environment. Much of the tooling for model-building comes from autonomous vehicle companies; Foxglove, for example, was founded by former employees at Cruise, General Motor’s erstwhile self-driving effort. And now those car companies are increasingly betting that their investments in ML tooling will allow them to compete with dedicated humanoid makers. Tesla is already trying this with its Optimus robot, and now both AV-focused Wayve and ride-share giant Uber have now launched robotics labs focused on humanoid form factors as R&D efforts. “I think you need to start in vehicles…manipulation robotics is like self-driving five years ago,” Alex Kendall, the CEO of Wayve, told TechCrunch. “The data infrastructure, the simulation, ML ops infrastructure, will probably be shared, but the specific world model for the simulator will be a different post-training. There’s going to be a lot of more more commonality than not, but then there’s going to need to be some some differences for different embodiments.” Kendall argues that it’s too early to commit to any one hardware platform—advances in sensors and other components are coming quickly, and a truly general model should be more agnostic. Théophile Gervet, the CEO of Genesis AI, a vertically-integrated humanoid robotics company that raised a $105 million seed round this year, disagreed, telling TechCrunch “we’re too early in this wave for a brain strategy to work; our take us there’s lots of opportunities to co-design hardware and AI.” Gervet also touched on another hot topic in the sector: How specifically to focus your physical AI business. Robotics companies that are targeting specific tasks are getting their robots out in the field—Gritt is building solar farms, Agility is deploying robots in industrial settings, and Bedrock is operating excavators autonomously. Meanwhile, general-purpose humanoids aren’t getting out of the labs. “No customer cares about the general purpose robot that works at 80% success rate,” Gervet said of the dilemma. “We see a lot of other players go general, but there is no value provided because there’s no vertical focus. .. but then, if you’re building [for a narrow] vertical on top of GPT 2, you’re going to get crushed by the company building on GPT 4.” The temptation to get invest in a specific vertical, however, is tempting because it provides not just revenue but also real-world deployment data. While task-specifc data might not have enough diversity to push general purpose models forward, it is an important for making a robot that adds value. Bedrock CTO Kevin Peterson noted that his company was just starting with excavation as a way to understand the challenges of “manipulation in the wild,” but plans to develop an intelligence layer that stretches across a series of construction machines. Managing all that data is a challenge, especially because of the density of visual and lidar data. Foxglove announced a new product this week, built on top of an Nvidia’s Cosmos open weight world model, that allows engineers to search that data with sophisticated natural language queries to build out evaluations and simulations. The goal is faster triage and debugging so model builders can iterate faster. So what will be the fabled ChatGPT moment for physical AI that Sam Altman recently said is just a few years away? Kendall points out that the largest robot deployment in the world are still consumer vacuum bots. For him, a ChatGPT moment would be something that excites consumers, not investors, who seem to be plenty excited already. “One example of that would be when you get eyes-off autonomy for less than $1000 [worth of hardware] in a car,” Kendall says; not coincidentally, his company is licensing models to car makers in an effort to produce just that. And that business, which he sees as a multi-billion dollar opportunity, will allow them to build a truly general embodied AI model. For Gervet, the moment when physical AI becomes real is “manipulation that just works out of the box. You can talk to a robot in natural language and have it do any basic task for manipulation, like say pushing, pulling, closing a laptop, cleaning up a table, whatever you want to do, and it works to some level of reliability, let’s say 80% plus out of the box— that’s roughly your ChatGPT experience.” Adrian Macneil, Foxglove’s CEO, looks at the question a bit differently. “There will not be a ChatGPT moment for robotics,” he told TechCrunch. “The thing that made ChatGPT a moment in time was the distribution—they went from zero to like a million active users in like a week…distribution in the real world is way harder than that, right? I would be very excited for the Apple II moment in robotics or the IBM PC moment in robotics. When can I buy like a home robot that is gonna start doing some useful and fun stuff?”
View

Surprise: Z.ai is the AI lab behind the mysterious Ox Alpha model
Over the weekend, the nerds werebuzzing with speculationover which AI lab was behind Ox Alpha, the mysterious new open-weight AI model launched onto OpenRouter anonymously and already topping benchmarks andleaderboardsagainst the best models. As many had expected, Ox Alpha was spawned by GLM-maker Z.ai, according toBloomberg. Z.ai confirmed that Ox Alpha is the newest iteration of its GLM series, which Hugging Face famously used recently to defend itself against an attack from OpenAI agents. The company said it will release the weights for Ox Alpha on Wednesday, after which developers can build on top of it. The company describes Ox Alpha as “a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.” The release of Ox Alpha adds to theburgeoning threat of cheap, capable modelsfrom China that could take real market share away from expensive frontier model companies like OpenAI and Anthropic. Earlier this month, Z.ai released GLM-5.3, which rivals Anthropic’s Fable 5 on certain benchmarks. TechCrunch has reached out to Z.ai for comment.
View

Bill Gates wants to see a robot tax and ‘Human Reserved’ jobs to mitigate harms from AI
Bill Gates posteda long essayto his Gates Notes site today, showing just how much the Microsoft co-founder has been thinking about the social impacts of AI. Gates is mostly in the Responsible AI camp, arguing thatPacing The Frontier, the open letter published by AI employees pushing for an AI slowdown, would be a good idea, but expressing skepticism that it would be sustainable. He also expressed excitement about AI’s benefits to science and healthcare, but worried about the labor impacts — all pretty familiar stuff if you followAnthropic’s policy work. But there were a few genuinely new ideas that could change the conversation if they take off. For starters, he proposes a “robot tax”: Right now, if you’re an employer and you hire someone, you pay payroll taxes on their earnings. But if you buy a robot, you can usually write it off right away as a business expense. The tax system nudges you toward replacing people with machines.A tax would slow the rush away from human labor a little and raise money for retraining and a stronger safety net. He also proposes setting aside certain jobs as “Human Reserved,” essentially barring AI from being used in certain tasks. It’s an interesting idea, and one that would be easy for policy-makers to enact: We might set something aside as Human Reserved for economic reasons. For example, we may do it because allowing machines to take over a certain role will displace a large number of people who can’t easily change jobs. You can’t tell a 55-year-old who has worked in construction their whole career that they need to go work at an elder care facility and expect them to find it fulfilling.Sometimes the decision to make something Human Reserved will be driven by other factors. In health, for example, imagine a robot giving you the awful news that you have an incurable disease. There’s no technical reason why it couldn’t. Yet it shouldn’t.The Human Reserved domain will evolve over time — for example, we should consider setting aside some jobs now and phasing in AI slowly over years or decades with a commitment to preserve some jobs. We should also note that both ideas would put a pretty serious dent in the profits of the major labs, which may be why we haven’t heard much about them until now. In addition, there are still a lot of unanswered questions about who exactly would set these rules and what exactly the rules would say, if they were to be rolled out.
View

Ex-Meta scientists want to bring visual AI to the factory floor
AI is transforming everything around us but, thus far, it has largely remained contained to the digital realm. Increasingly, however, startups are looking to take it into the real world. Perceptron, a startup started by two former Meta research scientists, is one such company. Founded in November 2024, the firm develops frontier vision models that aim to help machines more competently interact with their physical environments. This week, the company launched its latest model, Isaac 0.5, which its creators say is designed to provide machines with the ability to “perceive, reason and act” in industrial settings. Specifically, the software is capable of helping vision-guided robots navigate complex environments like warehouses or factory floors. It also helps companies extract visual intelligence from videos recorded by those bots. Isaac 0.5 is also being released as an open-weight model, so its parameters and training materials can be inspected by anyone. The startup, which recently raised $21 million in a funding round led by Bessemer Venture Partners, was co-founded by Armen Aghajanyan and Akshat Shrivastava, who previously worked for Meta’s Fundamental AI Research (FAIR), the tech giant’s AI research division. The duo see their software as the future of industrial automated deployment. “Physical AI today forces a false choice: generalist foundation models that need multiple dedicated cloud GPUs for every instance, or narrow models that handle perception or control, but never both,” the company says. Aghajanyan and Shrivastava say their tool is unlike existing models in the space because it is general-purpose, meaning that it’s not built for one specific, repetitive task. Instead, they say, the model is designed to be flexible depending on the particular environment (or situation) it is in. In an interview, Shrivastava asked me to consider what goes into a simple physical process like organizing boxes: “Imagine there’s a robot being deployed to sort packages right now. What are the tasks it would need to do?” Such a relatively simple task indeed consists of many steps. A robot would first have to read the label on the package, do some spatial analysis to understand where the boxes are, and decide which one to pick up. If it’s picking up a series of boxes, it would have to plan which boxes to pick up and in which order. Perceptron’s software is designed to help robots find their way through each step of the process. To be clear, the industry already has software that can help machines do most of those tasks, but there are few programs that are designed to do it flexibly. Where does the data for this algorithmic alchemy come from? Models like Isaac 0.5 learn operational skills by ingesting gargantuan amounts of video training data. Perceptron says its new model was fed on a million hours of what is known as general video to teach its algorithm to identify particular settings, visuals, and scenarios. The company also relied heavily on what is known asego video— video captured, typically through a GoPro or a wearable camera, from the perspective of a person completing a physical task — as well asUMI video, which are similarly used to teach AI systems movements by recording repetitive human actions. While Perceptron isn’t disclosing the sources of its training data, Shrivastava said that the company had “internally built petabyte-scale datasets that span across modalities, whether it’s images, text, video, etc. all the way through robotic trajectories.” The utility of a software that can help robots operate competently in warehouses is obviously vast, and Perceptron thinks it’s well-positioned to lead that wave of automation. The startup is ready to market its software to a variety of vendors, and thus potentially see its intelligence layer integrated into a broad array of industries. Those industries include manufacturing, logistics and warehousing, security, mobility, as well as media and entertainment. “Nothing like this really exists out there,” said Aghajanyan. “We’re really excited about it.”
View

Radar makes podcasts searchable — and usable by AI agents
Particle, theAI newsreader startupfounded by former Twitter engineers, is shifting its focus to a potentially more lucrative idea: indexing the spoken conversations buried in podcasts and making them discoverable. On Wednesday, the company introducedRadar, a podcast search engine that not only transcribes podcast audio but also understands what it means, enabling it to pull out key quotes and highlights. The solution has business potential, as it’s already attracted interest from hedge funds looking for data that their agents can’t see, explains Particle co-founder and CEO Sara Beykpour. “Hedge funds have been the highest-volume customers that are directly integrating with the API,” Beykpour told TechCrunch. While journalists and researchers could also make use of the tools, other top-paying customers have included AI search platforms and data resellers. (The search API provider for AI agents,Exa, for instance, is among Radar’s partners.) The idea itself stemmed from one of the Particle news-reading app’s most beloved features. The app had usedan APItosource interesting podcast clipsthat it then included alongside related news stories in the app’s feed. Particle’s team realized the product’s value, but also that it was somewhat trapped in the news reader. As the movement around AI agents began to gain steam, the company decided to pivot and focus on building an API for its podcast intelligence product. “Our vision is really to have all new media intelligence and all audio intelligence in that API. One of the reasons why it’s an interesting space is that most API agents and services crawl the web and they’re focused on text. We are providing that layer with audio,” Beykpour said. “Agents are generally blind to audio; they can’t see it unless something or someone has transcribed it.” With Radar, the company transcribes more than 130,000 podcasts, making it the largest transcribed podcast service in existence. This includes all the Apple Top 200 podcasts across its 135 verticals, with 20,000 episodes added to Radar’s index daily. The podcast transcriptions include speaker labels and rich metadata, as Radar understands the entities — people, companies, brands, products, and topics — being discussed. It’s also able to track mentions of these entities across podcasts and send alerts whenever they come up, either when the mention occurs or as a daily or weekly digest. The alerts, which can be delivered via email, Slack, or webhook, can be customized with filters. These let users configure Radar to only send alerts when certain guests appear and discuss a particular topic, for instance. The search can also be narrowed in other ways such as limiting it to top podcasts only. Radar can extract relevant self-contained clips, with timestamps, allowing users to both listen to and read the comments made. “We’ve pre-chosen notable clips, so if you can’t listen to the whole podcast and you don’t want to read a summary, this is the best way to just get an idea of what’s happening in that podcast,” Beykpour noted. Radar can also track the topics mentioned in the podcast, who or what was mentioned and when, listener ratings and reviews, the episode’s ads, and more. There’s even a dedicated podcast ads search engine that can find every episode where a given company advertises and track how it trends over time. This feature has additional monetization potential, alongside other tools offering political bias analysis, chart rankings data, audience size estimates, sponsorship data, and brand suitability. While all of this is available through Radar’s web interface,its real product is the API and MCP, which allows AI agents and other businesses to tap into this same intelligence programmatically. Radar is priced at $29 a month per seat, with a $399-per-month plan for businesses that includes 20 seats. API users have custom pricing, based on their needs. In the future, Radar plans to expand the service beyond podcasts to support other forms of audio, such as YouTube videos and news clips.
View
