
Hi everyone,
I was in London with the family last week and took last weekend off. It was a wonderful trip. London has quickly become one of my favorite cities. Between all the walking and its excellent public transportation, we covered a lot of ground.

Today's newsletter covers two weeks of research. Three patterns stood out: company memory is becoming AI infrastructure; the model is becoming one decision among many, rather than the main decision; and generating answers is getting much easier than judging whether those answers are good enough to use. All three move the difficult work into the controls around the model.
That last one is the subject of this week's essay. The practical question is whether the value of work good enough to use exceeds the cost of generation, review, correction, and failure.
Cheers.
Reza
📶 Patterns & Signals
Your agents' memory is becoming company infrastructure
AI tools are beginning to share company goals, decisions, and prior work. Asana is giving agents access to its record of work, while Tencent is testing shared team memory. Shared context reduces the need to rebuild company background for each task, but it also gives errors a longer life. An old forecast assumption, customer fact, or product claim can spread into every new task. Agent memory should not become a second system of record. When memory conflicts with a governed application, companies need a rule for which system wins, who corrects it, and how every agent receives the correction. Shared memory needs sources, owners, access rules, expiry, correction history, and a way to export or delete it - vs a dumping ground of all matters of records like meeting transcripts and slack direct messages.
Model choice is becoming a workload decision
OpenAI cut the price of its lower-cost GPT-5.6 models and now describes selection in terms of cost per finished task. Microsoft built smaller models for GitHub Copilot and Excel that it says handle common work more cheaply. Moonshot AI's Kimi K3 gives companies a model they can run themselves, along with the hosting, security, and support burden. A lower-cost model may be enough to classify incoming requests; a complex contract review may justify a stronger one. Compare quality, speed, privacy, and total cost for work good enough to use, including review and rework. Each production workflow also needs an owner, repeatable tests, a fallback, and a trigger for changing models.
Signals to watch
Business software is splitting into durable systems and replaceable features. AI can reproduce more features, but systems that hold trusted customer, financial, and operating data are harder to replace than thin applications with an AI button attached.
AI revenue is growing, but infrastructure spending is growing faster. Buyers should watch when the gap appears in contract renewals, usage tiers, or limits on capacity that had been generously bundled.
Large employers are hiring junior workers again. Watch whether the reversal lasts; workforce plans based on assumed agent productivity may have moved ahead of measured results.
MCP is starting to look like enterprise infrastructure. MCP is a common way for AI agents to connect to systems such as a CRM or ticketing platform. Its latest release strengthens security and normally gives companies at least a year to migrate away from a feature before it is removed. That makes agent integrations easier to plan and maintain.
🎯 AI can produce work quickly. Deciding what is good enough to use is harder.

After a year of tinkering with AI, building things, and talking with business leaders and practitioners, I have learned that AI can produce a lot of work and automate processes very quickly, at least in a pilot or demo. The difficult part is deciding which workflows are safe and useful enough to put into production. When a system can produce thousands of answers, a vague definition of “done,” what must be true before an output can be trusted and used, becomes a much bigger problem.
AI output still needs to be tested
In a study published in Science this month, researchers at the Arc Institute and Stanford used AI models to generate complete genomes for bacteriophages, viruses that infect bacteria rather than people. They selected 285 designs for laboratory testing; 16 produced viable viruses. The AI produced possibilities. Laboratory testing established which ones worked.
Most business work is less dramatic and not as headline grabbing as the news above, but Built Technologies offers a useful example. Its AI extracts information from real-estate finance documents, then a second AI step compares each answer with the source and sends uncertain cases to a person. Built and AWS say this reduced work from days to minutes. The important part is how the workflow defines the evidence required and who handles exceptions. Because the second check also uses AI, its decisions still need to be compared with known errors and monitored for what slips through.
Define what “done” means
The definition of done already sits in process guides, approval policies, required fields, brand standards, and the judgment of experienced people. AI makes more of it testable. “Every amount must match a cited clause” can become a check. “Non-standard payment terms require finance approval” can become a routing rule. Some judgment will remain subjective, so teams will still need to sample accepted work and measure downstream results, not just whether the artifact passed its checks.
How carefully you check AI output should depend on what the system is allowed to do. If it only drafts brainstorming notes for a person to rewrite, occasional mistakes are manageable. If it can approve a payment or send advice directly to a customer, the business needs clear tests, evidence for each decision, a named human approver, and a safe way to stop or reverse the action. More authority requires stronger proof that the system is working.
Defining “done” only works if the definition is right. In July, OpenAI reviewed 731 tasks in SWE-Bench Pro, a benchmark where AI coding agents receive a bug report, change real software, and pass automated tests to prove the fix works. OpenAI found that roughly a third of the tasks were flawed. Some bug reports left out necessary requirements, some tests rejected valid solutions, and others allowed incomplete fixes to pass. The lesson for business workflows is simple: a clean score can still approve bad work when the test itself is poorly designed.
Review time also belongs in the return-on-investment calculation. In a small Mayo Clinic pilot involving five emergency physicians and 710 visits, doctors using an AI scribe spent 4.3 minutes per adult visit working on the clinical note, compared with 1.8 minutes when using a human scribe. Physicians also wrote about 60% of the final note themselves with AI, versus about 31% with a human scribe. The study did not separate reading time from correction time, but it shows why the time people spend reviewing and completing AI-generated work belongs in the business case.
I know what you may be thinking: I just need to automate something quickly and show results. All this definition and testing will slow me down. It will, a little. But a fast pilot can prove that AI produces something; it cannot prove that the business should rely on it. This does not require a large governance program. It requires answering a few questions before you scale.
Before you scale
I would ask four questions before expanding an AI workflow:
Who owns the business result and has the authority to stop the workflow?
What will business success look like after rollout, and how will we measure it?
If we designed the process today, which existing steps or outputs would we eliminate rather than automate?
What evidence must accompany each AI output before it can be used?
The business outcome may remain, while today's reports, approvals, and handoffs may not. Automating an unnecessary output only makes waste happen faster.
Evidence may come from a fixed rule, a scoring rubric, or human judgment. The method should match the consequences of getting it wrong. Count the time spent checking, correcting, and recovering from errors. If the automated process costs more or performs worse than the current one, it is not a saving.
💬 Interstitials / Overheard

With OpenAi and Anthropic being in the news for their models breaking out of their labs and hacking other companies, this seemed appropriate for this week. The irony of this meme was that one week later, Meta disclosed that one of its own models had reached the internet during a security test and breached another company's system. A misconfigured test environment had given it internet access. Apparently the product request was fulfilled!
🌍 Meanwhile...

Google DeepMind's new WeatherNext model can predict a tropical cyclone's track, strength, and wind field better than previous systems. Its three-day forecast is about as accurate as the old two-day forecast, giving forecasters roughly one more day to prepare. The model can generate an ensemble of up to 1,000 forecasts, helping experts estimate uncertainty and unlikely outcomes instead of relying on one predicted path. The underlying research was published in Nature, and DeepMind is releasing the code and model weights so researchers can test and improve it. The gain will depend on whether emergency managers can turn the extra day into earlier evacuations and better placement of people and supplies.
📚 What I'm consuming
Companies Need to Stop Kidding Themselves About A.I., by former Lululemon CIO Julie Averill. She calls unsupported hope “AI wishing” and overstated progress “AI washing.” The distinction helps separate deployed capability, actual use, and a verified business result.
How an OpenAI engineer uses Codex and ChatGPT Work, from How I AI. ChatGPT Work holds the research and plan while Codex, OpenAI's task-running coding agent, executes.
The AI-Powered Transformation Office, from BCG. The proposed team coordinates AI adoption and controls. Its strongest rule is “minimum viable governance”: decide what AI may draft, what requires review, and what must remain a human decision before adding tools.
What Is AI Model Collapse? Why AI Could Forget Reality, from IBM Technology. A clear explanation of how repeated training on AI-generated material can erase rare knowledge while leaving a model fluent and confident. The practical lesson is boring and important: preserve human data, track where training material came from, and keep models connected to external sources.
The Hottest New AI Chatbot Is Just a Guy Answering Your Questions, by Caroline Haskins in WIRED. Tucker Bryant built ChatTJB, where “AI” means “average individual” and every response comes from him or one of ten volunteers. The joke is good, but the questions people ask are better: many do not need expertise; they want another person to think alongside.
🎙️ On Digital Edge

I joined Dominik Muggli on The Digital Edge Podcast to discuss whether companies should still buy most of their software now that AI makes building easier. Many thanks to Dominik for having me on the show. We talked about which applications are realistic build candidates, why the first version is usually the easy part, and how security, rollout, maintenance, ownership, and return on investment determine whether a custom tool survives. We also discussed why AI should not automate a broken process and why serious experimentation needs dedicated capacity.
🌙 After Hours
The Manchurian Candidate
Richard Condon | 320 pages | ★★★★★

It is a classic that has been adapted for film twice. The plot may seem outlandish, but its point is recognizable: politics is often shaped behind the scenes by people trying to control what the public thinks, rather than by what the public actually needs. I liked how quickly The Manchurian Candidate gets into the story and how little space it wastes. The book contrasts literal brainwashing with quieter manipulation inside families and politics. That makes it easier to suspend disbelief and think about what brainwashing means in today's social media maelstrom.
The ending feels rushed, and one late scene is strange even for this book. Still, the premise is original, the book moves, and the central idea remains unsettling almost seventy years later. A clear five for me.
🎙️ Listen
Prefer to listen? Quanta Bits is also available on Apple Podcasts and Spotify.
How this gets made
I collaborate with Spock, my AI agent. He researches extensively: scanning, filtering, and surfacing what is relevant across my business. I read, listen, and watch what resonates, and decide what matters. I provide direction, we draft together. The editorial judgment is mine. He would tell you the same. Most logical. 🖖