Another Daily AI Newsletter - July 24
⭐ Top Story: ChatGPT turns voice into a desktop control layer
OpenAI put ChatGPT Voice inside its desktop app, giving people a hands-free way to control their computer and direct multiple agents running in ChatGPT Work or Codex. The feature is rolling out globally on macOS and Windows for Plus, Pro, Business, Edu, and Enterprise plans.
9to5Mac reports that a user can start a task in voice mode, then ask ChatGPT to launch, check, or redirect work in other threads. Voice can use Computer Use, local files, and plugins. On macOS, Appshots can reference the active window, while ChatGPT Remote on iOS can monitor work running on another computer.
GPT-Live listens and speaks at the same time, making the conversation feel less rigid. When a request needs search, deeper reasoning, or a longer agent task, it can delegate the work to a frontier model in the background and keep talking. OpenAI says more than 150 million people already use ChatGPT Voice or Dictation each week.
That changes the role of voice. A person could ask Codex to inspect a codebase, send ChatGPT Work to revise a presentation, and request progress updates while both tasks continue. OpenAI voice product lead Atty Eleti told TechCrunch that the company sees voice becoming a primary interface for complex computing and long-running agent work.
The release is limited to paid desktop plans, and GPT-Live does not yet support video or screen sharing. Voice control also makes permissions and error recovery more important because a misunderstood instruction can change a file or send an agent in the wrong direction. OpenAI now has the pieces for people to manage AI work without staying at the keyboard.
AI adoption is broadening faster than full automation
Google’s ATLAS study examined 15 million AI interactions across more than 150 countries. Google found work-related AI use across 68% of occupations representing 90% of U.S. employment, but the typical occupation used AI for about 21% of its tasks. Fewer than 10% of work interactions fully automated a task.
The study also found that more than 86% of AI interactions occurred outside work. Manual and technical workers were using multimodal systems for diagnostics and troubleshooting, suggesting that adoption is spreading beyond office writing and software development.
Google says Gemini has passed 950 million monthly active users. The figure rose from 750 million in February and tripled year over year. TechCrunch reports that Gemini reached a 27.7% share of the consumer chatbot market in the first half of 2026, while ChatGPT’s share fell below 50%.
ChatGPT Health can now connect Apple Health and supported medical records for eligible U.S. adults. OpenAI says more than 300 million people use ChatGPT for health-related questions each week. The company says connected health data is not used for training or advertising, and warns that ChatGPT does not replace medical care.
The numbers describe broad adoption without broad autonomy. AI is touching many jobs and personal decisions, while most tasks still involve a person choosing the goal, reviewing the result, and carrying the work forward.
Software is becoming programmable by agents
Andrew Ng released OpenWorker, an open-source desktop coworker that runs locally. It supports models from OpenAI, Anthropic, Google, open-weight providers, and Ollama, connects to more than 25 services, and can run scheduled work. OpenWorker asks before consequential actions and currently offers a signed macOS build, while its Windows builds are not yet code-signed.
Notion introduced Notion as Code in beta. Teams can define workspaces, databases, and custom agents in TypeScript, deploy the configuration through an API, and track changes in Git. The release treats an operating workspace more like software that can be reproduced and reviewed.
Grok’s Build Workflows can coordinate hundreds of agents into one report. SpaceXAI positions the feature for large code changes, issue triage, and other work that can be divided across many agents before one system reconciles the results.
These products expose more of the agent system itself. Teams can define the workspace, select the models, schedule the work, approve sensitive actions, and preserve the configuration. Agent behavior is moving from an improvised chat into something closer to managed software.
AI products are becoming portfolios of specialized models
Runway launched a Media Router that selects an image, video, or audio model for each request. The router weighs quality, speed, and cost across Runway and third-party models available through Runway Dev.
Microsoft expanded its in-house MAI model family with image and voice releases. MAI-Image now powers Bing Image Creator and appears in PowerPoint and OneDrive. Microsoft says MAI-Image-2 reduced GPU cost by as much as 84% compared with GPT-Image-2, while MAI-Voice-2-Flash reduced GPU cost by as much as 89% in its contact-center workloads. Those efficiency figures are Microsoft benchmarks.
Sakana AI introduced Fugu Ultra as a multi-agent system delivered through one model endpoint. It dynamically selects and switches among a pool of models, while giving customers controls to exclude providers or models for compliance. Sakana’s benchmark comparisons are company-reported and still need independent evaluation.
Black Forest Labs previewed FLUX 3 as one model for images, video, audio, and action prediction. The early release points toward creative systems that share context across formats instead of handing each step to an isolated generator.
The model name is becoming less visible to the end user. Products are assembling portfolios and routing each request to the system that best fits the task, cost, latency, or compliance requirement.
AI compute is getting more specialized, from full racks to lunar rovers
AI chip startup Etched raised $300 million at a $10.3 billion valuation. The company says it has booked $1 billion in orders and begun testing its first systems. It is developing chips and cluster-scale memory around transformer inference, but mass production remains the largest execution risk.
AMD unveiled Helios, a rack-scale AI system scheduled to ship later in 2026. AMD says OpenAI, Meta, Oracle, Anthropic, and Microsoft plan to deploy the system. The announcement gives large customers another integrated option beyond NVIDIA’s rack architecture.
NVIDIA is sending a Jetson GPU toward the moon. Lunar Outpost plans to use the module to process lidar and control its next rover locally. The launch is expected before the end of 2026 and would likely place the first GPU on the lunar surface.
The hardware spans three distinct constraints: dense data-center throughput, lower-cost inference, and autonomous decisions where cloud connectivity is unavailable. Specialized AI is being designed around where the computation must happen.
Quick Hits
Claude Voice now uses Opus, Sonnet, or Haiku and can reach connected work apps. Claude can access Gmail, Google Calendar, Slack, Canva, and Notion during a conversation. Free users are limited to Haiku and one connected app.
Qwen released Audio 3.0 TTS with 16 languages and 20 Chinese dialect regions. The system supports up to three minutes of speech, natural-language delivery controls, and 86 inline expression tags. Its leaderboard results are Qwen’s own.
LangChain says Salesforce has accepted 100 million lines of AI-generated code. The figure comes from LangChain’s announcement and offers one enterprise-scale measure of agent-assisted development.
ChatGPT Sites added basic performance analytics in public testing. Builders can inspect early traffic and performance signals without leaving the product.
💰 Funding & Moves
Sierra acquired Takeoff, a three-person startup building long-horizon agents. Takeoff says it grew from zero annual recurring revenue to nearly eight figures since the beginning of 2026. Its Horizon platform will become part of Sierra’s work on outcome-based, industry-specific agents.
Cognition acquired The Interaction Company, the team behind the text-message agent Poke. Cognition says Poke handled 100 million messages in three months and reached hundreds of thousands of users. The product will continue while the team applies Cognition’s models and infrastructure to personal agents.
🔬 Research Radar
RedNote says its dots-note-3.0 model received a perfect 42 out of 42 on the six IMO 2026 problems. The internal beta model produced complete natural-language proofs and received full marks on every problem. Only seven of the 666 human contestants also earned a perfect score.
RedNote says it plans to open-source the model, but has not yet released its weights or a complete account of the prompts, tools, sampling, and retry budget behind the result. The score is a significant reported reasoning result, not yet an independently reproducible benchmark.
U.S. and U.K. evaluators found that Kimi K3 can assist with offensive cyber work, but trails leading closed models. Kimi K3 scored 32% on the 41-task ExploitBench, ahead of GLM-5.2 at 24%, but achieved arbitrary code execution on none of the tasks. In a simulated 32-step corporate-network attack, it reached step 17 on average and completed the full path once in ten attempts. The evaluators describe the results as preliminary and note that Kimi was tested on a narrower set of benchmarks.
Researchers interviewed by TechCrunch also challenged a White House claim that Kimi’s strength came primarily from distilling Anthropic’s Fable. Their argument is that Kimi’s release arrived too quickly after Fable for large-scale Fable distillation to explain the model, and that advanced reinforcement learning would require substantial time and infrastructure. Moonshot has not publicly detailed Kimi K3’s complete training process.
🛠️ For Builders
Trello launched an official MCP server for boards, cards, and lists. The first release connects one workspace and excludes destructive delete actions.
Hugging Face added native Nunchaku 4-bit diffusion loading to Diffusers. Its published tests show up to 50% less VRAM and roughly 30% lower latency, with results varying by model and hardware.
NVIDIA published a hosted reinforcement-learning tutorial for Nemotron 3 Nano. In its small 32-problem math experiment, the score rose from 21.9% to 90.6% for less than $5.
📘 AI Term of the Day: Manager agent
Google defines a manager agent as an agent that controls one or more sub-agents.
The new ChatGPT Voice experience is a practical example. The manager interprets the spoken request, starts or redirects specialized agents, checks their progress, and returns a combined update. That hierarchy can make parallel work easier to control, as long as permissions and responsibility remain visible.
Google’s definition | Go deeper with Google’s architecture for multi-agent systems


