EPISODE · Apr 15, 2026 · 29 MIN
Proof that Opus 4.6 Is Getting Worse, Ramp AI Coworker, MiniMax M2.7 & More (This Week In AI)
from Agents Hour · host Mastra
Mounting evidence that Claude Opus 4.6 has been degraded — BridgeBench shows a 15-point accuracy drop on their hallucination benchmark, and AMD's Senior AI Director found median thinking collapsed from ~2,200 to ~600 characters between January and March. The hosts share their own experiences, and they line up. Meanwhile, a claim surfaced that Cursor Agent is a rebranded version of Claude Code, running behind a local proxy with a find-and-replace engine that swaps "Claude" for "Cursor" in system prompts. Cursor's Michael Truell responded, saying it was a sub-1% A/B test. The hosts break down both sides. On the shipping front, Anthropic launched Claude Managed Agents in public beta, released Claude for Word, shared details on Claude Mythos Preview — including speculation that it's a looped language model based on a ByteDance paper — and expanded its Google/Broadcom partnership for multiple gigawatts of compute. Their run rate reportedly jumped from ~$9B to $30B in four months. Sam Altman published a personal blog post revealing that someone threw a Molotov cocktail at his house. Plus: why senior executives are voluntarily dropping title to join AI companies, Ramp's internal AI productivity suite Glass, Ramp Labs' Latent Briefing paper showing 31% token savings for multi-agent systems, Scale AI's Muse Spark model now powering Meta AI, GLM-5.1 breaking into Code Arena's top 3, MiniMax shipping MMX CLI and open-sourcing M2.7, and widespread benchmark cheating exposed across nine agent benchmarks. 🔗 LINKS https://x.com/bridgemindai/status/2043321284113670594 https://x.com/hesamation/status/2042979500103815306 https://x.com/steipete/status/2042615534567457102 https://x.com/claudeai/status/2041927687460024721 https://x.com/claudeai/status/2042670341915295865 https://x.com/alexalbert__/status/2041579938537775160 https://x.com/ChrisHayduk/status/2042711699413926262 https://www.anthropic.com/news/google-broadcom-partnership-compute https://x.com/noahzweben/status/2042332268450963774 https://blog.samaltman.com/2279512 https://x.com/aakashgupta/status/2042684298671853903 https://x.com/sebgoddijn/status/2042285915435937816 https://x.com/ramplabs/status/2042672773747589588 https://x.com/alexandr_wang/status/2041909376508985381 https://x.com/arena/status/2042611135434891592 https://x.com/minimax_ai/status/2042644651333816338 https://x.com/minimax_ai/status/2043132047397659000 https://x.com/adamlsteinl/status/2042655187613995026 AI Agents Hour is a weekly livestream hosted by Mastra CPO Shane Thomas and CTO Abhi Aiyer. Airing Mondays at 12PM Pacific on YouTube and X, the show covers breaking AI news, agent development techniques, and features interviews with industry experts building AI applications today. 📚 MASTRA RESOURCES Mastra: https://mastra.ai Mastra on X: https://x.com/mastra_ai Mastra Discord: https://mastra.ai/community/discord Mastra GitHub: https://github.com/mastra-ai Learn Mastra in the world's first MCP-Based Course: https://mastra.ai/course Principles of Building AI Agents (Book): https://mastra.ai/books/principles-of-building-ai-agents Patterns for Building AI Agents (New Book): https://mastra.ai/books/patterns-of-building-ai-agents MASTRA? Mastra is an open-source TypeScript framework designed for building and shipping AI-powered applications and agents with minimal friction. It supports the full lifecycle of agent development—from prototype to production. You can integrate it with frontend and backend stacks (e.g., React, Next.js, Node) or run agents as standalone services. If you’re a JavaScript or TypeScript developer looking to build an agentic or AI-powered product without starting from first principles, Mastra provides the scaffolding, tools, and integrations to accelerate that process. ⏱️ CHAPTERS 00:00 Intro 00:28 Is Opus 4.6 Nerfed? 04:00 OpenClaw vs Anthropic 04:41 Claude Managed Agents 07:09 Claude for Word 07:33 Claude Mythos Preview 09:51 Anthropic x Google/Broadcom — $30B Run Rate 10:38 Claude Code Monitor Tool 11:01 Is Cursor Just Claude Code? 12:49 Sam Altman's Personal Post 14:29 Executive Compression 17:04 Ramp Built Every Employee an AI Coworker 19:19 Latent Briefing 21:53 Is Meta Back in the Game? 24:46 GLM-5.1 Hits #3 in Code Arena 25:41 MiniMax MMX CLI 26:29 MiniMax M2.7 Open Source 27:28 Widespread Benchmark Cheating 29:07 Outro
Embed this episode
NOW PLAYING
Proof that Opus 4.6 Is Getting Worse, Ramp AI Coworker, MiniMax M2.7 & More (This Week In AI)
No transcript for this episode yet
Similar Episodes
No similar episodes found.
Similar Podcasts
No similar podcasts found.