Five ideas, papers, and products shaping our investment thinking, and what we think they mean. This series is powered by an AI assistant that helps synthesize recurring themes from our discussions, alongside our own reflections.
Today’s best AI models can control a robot arm, but they all failed at household chores. Researchers tested seven leading AI models at controlling robot arms across about 1,000 simulated tasks. The best model, GPT-6 Astra, finished 42% pick-and-place tasks. But on household chores, where the robot has to move several objects around a room in a sequence, every model failed every time. We think LLMs that decide what a robot should do are improving fast and will soon be cheap. The value is in getting the robot to actually carry it out reliably in the real world, i.e. control, sensing, and data from real deployments.
A new method taught an LLM-driven robot to fold towels and handle plates with 10 minutes of real-world practice. Learning by trial and error on real robots is slow and expensive. So researchers at Amazon FAR, UT Austin and CMU had a coding agent practice in a rough simulation first, then gave it five tries per task on a real robot. After each failure, the agent tested fixes in the simulator and kept only the ones that worked. It succeeded on 26 of 30 real attempts, while the best alternative managed 3. We think this is a real unlock: teaching robots new tasks normally takes hours of robot time and lots of real data, and this needed a rough sim and a few real tries.
Where robots already work reliably, in factories and warehouses, the problem isn’t capability. It’s cost. Anthropic’s economists rated 7,594 physical job tasks on whether a robot can do them today and whether it would cost less than a human. Robots can handle about three-quarters of them, almost entirely in controlled settings like factory lines and warehouses; in messy environments like homes, that drops to about 1%. Even where robots can do the work, they’re cheaper than a person for only 0.3% of tasks. We think the winners will be the companies that make the unit economics work, e.g. robots-as-a-service priced per task instead of hardware sold up front.
The bottleneck for AI inference chips is memory bandwidth, not compute. On No Priors, Fractile’s CEO argued that running large models is limited by how fast a chip can move data out of memory. Over 20 years, compute grew about a millionfold while memory bandwidth grew only ~40x. Fractile claims its custom DRAM design delivers ~25x the bandwidth of today’s HBM-based chips. We think memory architecture will increasingly decide the cost of inference. We’re looking for memory-first chip startups and for tools that shorten chip design cycles.
Anthropic let Claude agents trade on behalf of 201 employees, and their biggest problem was understanding what people wanted. In Project Swap, each agent interviewed its owner about their reading tastes, then traded used books with the other agents. On average, people ended up with about their fifth-favorite book out of ten. About 85% of the shortfall came from agents misjudging preferences, not from bad trading. We think agentic commerce is bottlenecked on understanding the user, not on payment rails. Stablecoins can handle settlement; the edge goes to whoever models customer intent best.
We’ll share another edition next week.







