AI NewsOpenAI is building AI agents for everything. Will everyone use them?
OpenAI is building AI agents for everything. Will everyone use them?
9:37 PM IST Ā· August 24, 2026

How much control are you willing to give an LLM over your digital life? Getting the most value from a model means giving it the keys. For a control freak or the AI-hesitant, it seems like a lot. For Andrew Ambrosino, the lead engineer for OpenAIās desktop app, itās the only way to test the future, which is why that app now has access to, and control over, his inbox, his Slack account, his phone, apps like Notion and Figma, and more. āIf Iām asking it to write a document, is there a possibility that itās going to pull from a private DM on that subject and not know that itās not supposed to share some info? Yes,ā Ambrosino told TechCrunch. āIāll do it for the job. I will take the personal hit here and there if I have to. And I havenāt had to.ā Ambrosino works on OpenAIās biggest bet, ChatGPT Work, which was released last month and is available on the companyās lowest subscription tier, for $20 a month. The product is intended to allow white-collar workers to field AI agents ā hooking LLMs up to the digital workflows used by accountants, investors, doctors, and everyone else whose day-to-day is dominated by their computer. OpenAIāsmarketing copyputs the goal succinctly: A world where āwhere [artificial] intelligence goes beyond answering questions to helping everyone turn their biggest ideas into reality.ā For software developers, that shift is already happening, but itās been slow to spread to other departments. ChatGPT Work is a modified version of the companyās Codex coding tool. Itās meant to give non-engineers a version of the same functionality that software engineers already get from agents: an AI tool that doesnāt just answer questions, but completes multistep projects on its own. āIn this new factor, ChatGPT can actually do entire, very complicated tasks for you all autonomously in a way that is delightful and safe,ā Thibault Sottiaux, who leads OpenAIās core product work, including Work, told TechCrunch. āItās the very mission of OpenAI ā to bring everyone along.ā Commercially, that matters a lot. Agents that work for longer stretches burn through more tokens, which makes them more lucrative for OpenAI on a per-user basis. Reaching new professions is crucial ā not just for OpenAI, but for the industry at large. If coding has proven lucrative territory for AI labs, itās still a tiny subset of the professional work AI tools need to enable if these companies are to justify their massive investment in training and computation. While labs have been focused on software engineers, vertical-specific competitors like Harvey (for law) and Clay (for sales) have been chasing those customers with a model-agnostic approach, meaning theyāll plug in whichever AI works best at the time. Industry analysts see this as one of the major challenges facing OpenAI and its competitors. āIf the labs cannot rapidly get ahold of the key complementary assets needed to scale AI in the market, value will accrue elsewhere,ā Christian Cataliniwroteon a16zās āItās time to buildā blog. Making the AI apps work for people who arenāt software engineers requires more hand-holding. OpenAIās non-engineering workforce, like the communications and finance teams, started using Codex āat a time that it was actively hostile to themāasking them about code and showing them, āoh, you have an empty diff for this thing,āā Ambrosino said, referring to a technical readout meant for software changes. āSo, we started to make it more general purpose between February and now.ā An OpenAI-backed study found that in June, 98% of OpenAI employees were using Codex, but just 17% of organizational subscribers and less than 1% of individual subscribers were using the agentic coding tool. That difference between near total adoption inside the company and negligible adoption outside it is the challenge and opportunity for the company.āThe more value and the more utility that we generate for users, the more they will be willing to also pay for some part of that utility, and thatās how weāve always seen ChatGPT as well,ā Sottiaux said. āYou sit there and youāre like, āof course I want to pay $20 bucks a month for this,ā because the value that you get is so much more.ā To understand that disconnect, it helps to understand what OpenAIās engineers are building. Every LLM requires what engineers call a āharnessā ā the software wrapped around a model that decides what information it sees, which tools it can use, and how it presents its answers back to you. If you want that model to do stuff ā to become an agent ā the harness gives it tools and instructions for using them on long-term tasks. For developers, a command-line interface (CLI) that enabled LLMs to code was enough to change the way software was built and deployed. But most people arenāt using CLIs; thereās a reason Windows replaced DOS. An agentic product that goes beyond software engineering is āgoing to be something that plays with the messy world of your life and your tools and websites that were built in 1995 and never updated,ā Ambrosino told TechCrunch, explaining that the experiences his team is building are vital to expanding access to useful AI. Consider apps like Claude Code and Codex: They unleashed āvibe codingā by abstracting away all the actual software writing, and letting users just tell the model what they want in a program. Now, OpenAI wants to make functionality found in tools like OpenClaw, which coders use to put LLMs to work, as easy as prompting. āWithout these products in front of the model, experts would know how to get the same results, but you wouldnāt get to a billion people using the thing,ā Ambrosino said. That trade-off between what power users need and what mainstream adoption requires plays out in internal debates at OpenAI, where some employees argue that a button is unnecessary if users can just ask the model directly. āWe push back on [that] ā because itās very early,ā Ambrosino said. āDiscoverability matters in this phase, and at some point we wonāt have the button.ā Work has a few more buttons for selecting projects and plug-ins, but it aims for the same magic box interface as other OpenAI products. He compares it to skeuomorphism, the fading practice of making digital tools look like the physical objects they replaced, like a calculator app made to look like a pocket calculator. āThat stuff wasnāt just cringe design. That actually helped get people into this [and] make the transition,ā Ambrosino said. OpenAI wouldnāt say how many people used Work versus Codex, but the joint app is used by just 20 million people, compared to more than a billion users the company says are prompting ChatGPT online. For now, OpenAI is pitching this tool as best suited for routine, data-intensive coordination tasks. Its employees are setting upweekly metrics reports, for example, and making spreadsheets intoplanning tools. Iāve spoken to VCs using agents to assemble relevant communications and analysis about companies into investment memos, and ops teams spinning up bespoke dashboards and data visualizations. Sam Altman is using it toplan his vacations. One OpenAI engineer described asking the program to look at a Slack conversation about an engineering problem and āmake some charts,ā then receiving back a series of insightful plots. āThere is a deluge of information for the average worker or employee of any of these companies, including myself,ā Akshay Nathan, who leads the product engineering team at OpenAI, said. āWeāre actually quite limited by our ability to parse everything thatās available to us, and then take action on it. That information lives in all these system records tools [like, Salesforce]ā¦the value of ChatGPT is you already have access to this, but now youtrulyhave access to it.ā This, then, could be the digital personal assistant that AI evangelists dream about. As with Claude Cowork or Perplexity AI browsing agent, ChatGPT Work links agents to your existing workspace ā email, web browser, a slew of SaaS platforms ā and puts that context to work for you. When the system works, it can be impressive: I asked ChatGPT Work to get my sonās weirdly-formatted preschool calendar out of my email and put it into my Google Calendar, and it did, saving me a lot of repetitive data entry. Hopefully now I wonāt forget the school potluck or fail to arrange vacation childcare. I didnāt trust OpenAI with access to my inbox, source interviews, or story drafts (fear not, AI haters) and wouldnāt let it have access to my bank account, but I believe it would be more useful had I the faith. I tasked it to do financial analysis on publicly traded companies that I cover, and it delivered an auto-updating dashboard of metrics for me; it made a queryable database of space launches, a task Iād previously had to accomplish by writing Python scripts. It also sends me a weekly email about new AI research posted at academic clearinghouses. Iāll keep experimenting with it. While asking the model for something is intuitive, giving it what it needs to take action isnāt as simple. Setting up the permissions for agents to access, say, a cloud drive was confusing and circular ā I tried multiple times to give it just āreadā access and received error messages. The model itself wasnāt too helpful, but eventually on the mobile app, a dialog box popped up to tell me that only complete access would make it work. Many important settings are only available on the web app, so I frequently found myself working in both at the same time. Sometimes ChatGPT Workās limitations are baffling ā link it to your Google calendar and it can create events, but not new calendars. And donāt bother trying to do anything unless the effort level is high, otherwise youāve got the worst intern youāve ever worked with. Thatās common advice from AI early adopters, who fear that frustrated newbies will give up. Joe Gershenson, the engineering lead for OpenAIās harness, admitted that effort settings arenāt intuitive for new users yet ā āthere are things that we can do better to help them get the right level of reasoningā¦ā he said, adding, āWatch this space.ā OpenAI faces another important challenge breaking into normie white-collar work: Most workflows arenāt as measurable ā or evaluable ā as code. Software either works or it doesnāt, and even that distinction reduces the nuance about what makes code good or bad. A good presentation, business strategy, or sales pitch isnāt as easy to evaluate or trace. āOne of the unique challenges with a product like this is just that it can really do anything,ā Ambrosino said. When I asked which specific problems the team designs around, and which workflows it targets, the engineers I spoke with demurred, saying that was a question for OpenAIās research team. OpenAI later provided an answer, telling TechCrunch that it uses its benchmarkGDPval, drawn from 44 occupations and hundreds of knowledge work tests, and supplements that with user feedback. A less official answer is that it comes from OpenAI employees themselves. Said Ambrosino, āWe have to always parse out ⦠are we doing the workflow that everybody else will be doing, or are we weird?ā The early adopters of the app itself will create valuable traces with their actual usage ā much of the success of coding tools is built on similar data collection ā assuming they donāt opt out of making it available for training. (I did). Despite all the attention on the model interface, OpenAIās engineers were reluctant to answer a fairly simple question: What sets Codex and ChatGPT Work apart from Claude Cowork, or other competing agentic harnesses intended for a mass user base?āItās going to be a really disappointing answer for you, and Iām sorry, but the honest answer is that I really donāt look at the harnesses that theyāre building,ā Gershenson said in a typical answer. āThe Mad MenāI donāt think about you at allāmeme comes to mind here.ā Frankly, I donāt believe them, if only based on the extreme similarities between the productsā user interfaces, the need for competitive intelligence at any business, and because the first thing ChatGPT Work asked me to do when I started it up was port over my Claude Cowork data. Itās understandable if Claude Code is a sensitive topic around the OpenAI offices. Their corporate rival defined the market for AI coding and launched a revolution in how software engineers do their jobs. Itās additionally frustrating because OpenAI had the idea first, but didnāt quite harness it correctly. When OpenAI first developed Codex as a web app, the engineers got a bit over their skis ā or, as Ambrosino puts it, āa bit more AGI-pilled.ā In short, they bet on the model being smart enough to handle a task entirely on its own, with minimal user input. Built shortly afterward, Claude Code was oriented around a back-and-forth conversation with the user. If you gave it a problem, it would survey the possibilities and give you three or four options for proceeding. Once you chose, it would go a little further and then check back again, giving continual updates and leaving less room for the model and harness to make mistakes. Anthropicās approach proved more effective, even if it demanded more work from users. ā[Our] product was a little ahead of where the model and harness was at the time,ā Ambrosino says now. OpenAI eventually followed suit by adding more opportunities for users to interact with the model. That became the Codex we know today, with desktop and mobile apps. Using download statistics as a proxy for interest in the programs, Claude Code was more in demand until April of this year, but now Codex has taken a slight lead;surveysof enterprise use also suggest OpenAI is catching up. Part of that lead is getting the product-market fit right, and part comes from complaints about safety restrictions on Anthropicās models and compute shortages. OpenAIās steps toward more human-centric harness continue with ChatGPT Work, but the engineers I spoke to insisted the key differentiator is the strength of OpenAIās latest powerful and cost-effective models. āThe frustrating answer is that a lot of times it is the model, and one thing that we have tried to do really well with this app is fully leverage the model,ā Ambrosino said. That explanation returns to the ābitter lessonā learned by AI researchers that a better general model is more important than specific domain experience. For true believers, the harness is a temporary crutch, not the moat. āYou could get good results in the short term by adding a whole bunch of extras ā if and thens and tools ā but like, come on, the next model is going to come out in a couple of months and make that obsolete,ā Gershenson told TechCrunch. His team focuses on the simplest ways to expose the model to the tools and context it needs ā and no more. āThe goal of good harness engineering is to⦠be more precise about what information the model really needs to solve your problem, because the models are getting better and better at doing that if you simply let them do their thing,ā Gershenson said. There is an open question, though, if most people or models are ready for that. Ethan Mollick, the Wharton School of Business professor who studies AI tools in the workplace, still sees Claude as more user-friendly,writingthat āChatGPT tends to want to do magic & just do it for you, while Claude does comparisons & shows them, repeatedly asking for input & feedback and doing A & B tests.ā Sottiaux, and perhaps OpenAI at large, disagree, arguing that the conversational nature of the app is better than learning how to use an application. āWe definitely see that the world seems to be ready,ā he says of the app. āThis is why weāve had incredible adoption.ā Still, itās not clear that a model-specific harness is even the right bet for maximizing a model. Comparisons run by companies likeComposioandDatabricksshow that different harness and model combinations deliver different performance on coding benchmarks. Databricks found that Pi, an open source harness published by the software company Earendil, outperformed Codex while using the same GPT 5.5 model. Pi has been used to build projects like OpenClaw and Cloudflare OS. Piās creator, Mario Zechner, says his intentionally minimalist harness is evidence that an AGI-pilled approach can work, at least for software engineers and coding tasks. What it lacks in explicit features, he says, is made up for by its ability to modify itself and build its own interfaces. He sympathizes with the challenge that OpenAIās engineers face in expanding their user base beyond engineers. āEverything is coding agent shapedā¦the reason is that they only have training data for coding agent tasks,ā he told TechCrunch. āSay Iām in management, I make a decision today, and the outcome happens months later. You cannot capture that in a simple trace of a user and agent back and forth, so all of these kinds of tasks and anything that you donāt digitize is inaccessible to a model to learn.ā Like other open source providers, he sees the big labās effort to push their harnesses as a way to lock-in users; āThey need to own the entire stack; otherwise, they just become a model provider and then need to compete with Chinese models.ā He and other engineers TechCrunch spoke to felt that the insight into token spend and agent behavior in frontier labsā harnesses is too limited. In a sense, thatās less meaningful to non-technical workers, but as with the coding tools, uptake at the scale OpenAI hopes for will eventually force harder conversations about cost. For example, messing around on a $20-a-month subscription, I used more than 80 million tokens in four days, which cost $65, according to the modelās analysis (thereās no dashboard in the app). Thatās a subsidy of more than 3x the subscription price for four days of casual use alone. āWe are working every day to push the frontier on efficiency,ā Sottiaux said, pointing to a recent 80% price cut for users of OpenAIās Luna model. āIf you wake up six months from now, you should be able to do all of the same [tasks] with less spend.ā The other relevant question is whether these apps create a lock-in effect on customers through data retention, or the sheer pain of configuring access to all the plug-ins and their permissions. Inside OpenAIās wood-paneled, plant-filled headquarters, which I visited in July, the atmosphere was calm but slightly tense; these are people with a lot to do. The engineers I spoke with were constantly monitoring their laptops as we talked, and rushing from meeting room to meeting room. Nathan, the head of the product engineering team, said the focus remains on āthe promise of the magic box, but I still think thereās too much complexityā¦Iām very optimistic that we can solve it, with the model and in a truly AI-native way.ā
read more



