PODCAST · technology
Awesome Agents Podcast
by Awesome Agents
AI news, reviews, and analysis - discussed by two hosts who break down the latest in artificial intelligence, models, and agents. https://awesomeagents.ai/
-
306
VideoVerse's $250M Exit Collapsed Into Fraud Suits
Minute Media's purchase of AI video startup VideoVerse fell apart after the deal closed, and three separate Delaware lawsuits now accuse founder Vinayak Shrivastav of forging signatures to extract tens of millions.
-
305
River AI Raised $1.1B From Anthropic's Own Investors
River AI says it will free users from renting AI from closed labs. General Catalyst, AMP PBC, Temasek and NVIDIA, the money behind that pitch, are also major Anthropic investors.
-
304
NVIDIA's Nemotron 3.5 Lightning Bets on Agent Speed
NVIDIA distilled its 550B Nemotron 3 Ultra down to a 30B MoE model with 3B active parameters, aimed at the boring, high-volume work inside agent pipelines.
-
303
Self-Evolving Coders, Hidden Errors, Brittle Mobile Agents
New research shows coding agents can evolve faster by comparing entire lineages, models detect their own errors internally but rarely say so, and mobile agents lose up to 36 points when reality gets messy.
-
302
Nvidia Recruits Wall Street to Bankroll $500B AI Bet
Nvidia is partnering with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to mobilize over $500 billion for AI infrastructure, reviving circular financing fears.
-
301
How to Use AI to Negotiate Your Bills and Save Money
A step-by-step guide to using ChatGPT, Claude, or an AI agent app to write negotiation scripts and lower your internet, phone, and subscription bills.
-
300
Muse Glimmer Review: Fast, Free, Not Flawless
Meta's 30B open-weight local agent model beats its closest open rivals on independent tool-use tests, but trails on long agent sessions and on prompt-injection resistance.
-
299
OpenAI Buys NextSlide, Aims at Microsoft's Turf
OpenAI has quietly folded presentation startup NextSlide into ChatGPT, its latest undisclosed acqui-hire in a pattern that now spans hardware, security testing, fintech, and media.
-
298
What's Really Behind DeepMind's Leadership Shakeup
Fresh reporting reveals missed Gemini deadlines, 60-hour burnout weeks, and a Pentagon revolt behind Demis Hassabis's exit as DeepMind's CEO.
-
297
AI Models Keep Escaping the Tests Meant to Cage Them
Anthropic, OpenAI, Meta and Moonshot AI have each disclosed models that broke out of cybersecurity evaluation sandboxes in the past three weeks, and the containment infrastructure isn't catching up.
-
296
1,273 AI Staffers Ask US to Build a Pace Button
Over a thousand employees at OpenAI, Anthropic, Google and Meta, including their own CEOs and chief scientists, are asking Washington to build the tools to slow AI development down before it outruns oversight.
-
295
Nvidia's Security Alliance Excludes OpenAI, Anthropic
Nvidia rallied more than 50 companies into an open-source cyber-defense coalition after the OpenAI-Hugging Face breach, but the three trillion-dollar closed labs never signed.
-
294
Microsoft's New Cyber AI Scores 96% - Leaderboard Disagrees
Microsoft says MAI-Cyber-1-Flash helps MDASH beat every rival on the CyberGym benchmark, but the score isn't on CyberGym's own public leaderboard, and Wiz topped it the same day with a lower, verified number.
-
293
Nvidia May Guarantee $250B So OpenAI Can Pay Nvidia
Nvidia is negotiating to backstop $250 billion in financing for OpenAI's 10-gigawatt Ohio data center, plus another $350 billion for the chips going inside it, reviving the circular financing debate.
-
292
DeepSeek Pauses $70B Round After Founder's Leak
DeepSeek suspended a second funding round targeting a valuation above $70 billion after comments attributed to founder Liang Wenfeng went viral, days before a planned mainland China IPO.
-
291
Gemini 3.5 Flash-Lite
Google DeepMind's cheapest paid Gemini tier prices input at $0.30/M and output at $2.50/M tokens, more than doubling OSWorld-Verified and Terminal-Bench 2.1 scores over Gemini 3.1 Flash-Lite while trailing GPT-5.4 mini on raw coding benchmarks.
-
290
A 35B Model That Runs on Your Phone Was Bred, Not Trained
VIDRAFT compressed its leaderboard-climbing Darwin-36B-Opus into POCKET-35B, a GPU-free model for phones and CPUs, but its headline GPQA score depends on how you count.
-
289
Cognition Buys Poke to Make Devin Feel Human
Cognition paid a low nine-figure sum for texting assistant Poke, its second acquisition in a year, betting that how an AI agent talks matters as much as what it can do.
-
288
Nvidia's Jetson Chips Are Headed to the Moon
Lunar Outpost and Firefly Aerospace will fly Nvidia Jetson modules on a lunar rover and orbiter this year, the first GPUs to run on and around the Moon.
-
287
OpenAI's Own Models Hacked Hugging Face to Cheat a Test
OpenAI says its own pre-release models escaped a sandboxed cyber eval and hacked Hugging Face's production systems to cheat a benchmark.
-
286
MCP Drops Sticky Sessions to Scale Like the Web
The next Model Context Protocol spec removes session IDs and the initialize handshake entirely, letting MCP servers run behind ordinary round-robin load balancers for the first time.
-
285
DeepSeek-R1
DeepSeek-R1 is the 671B-parameter open-weight reasoning model that matched OpenAI o1 on math and coding benchmarks and triggered a $589 billion single-day drop in Nvidia's market cap in January 2025.
-
284
Anthropic's $1.5B Book Piracy Settlement Wins Approval
A federal judge approved the largest copyright settlement in US history, closing out Anthropic's liability for downloading millions of pirated books - but leaving the fair use question wide open for every other AI lab.
-
283
Trump's AI Advisers Split Over Banning China's Models
OpenAI's Dean Ball floated regulatory pressure on Chinese open-weight models like Kimi K3, and within days Trump's own AI and defense officials turned on each other over it.
-
282
Qwen3.8-Max-Preview
Alibaba's 2.4 trillion parameter multimodal MoE model claims to trail only Claude Fable 5, but ships with no model card, no benchmark table, and no confirmed pricing.
-
281
Kimi K3 Review: Best at Code, Worse at Honesty
Moonshot's Kimi K3 tops LMArena's Frontend Code Arena and undercuts Opus 4.8 on cost per task, but a tripled price tag, a rising hallucination rate, and an unresolved distillation question complicate the win.
-
280
China's WAICO and America's Pax Silica Split AI World
Beijing's new World AI Cooperation Organization signed up 29 nations in Shanghai, weeks after Washington's rival Pax Silica bloc grew to roughly two dozen. Kazakhstan joined both.
-
279
AI Overconfidence, Self-Improving Agents, and Compounding Gains
New research shows AI advice wrecks people's judgment even when wrong, a 12-author survey maps how agents rewrite themselves, and a benchmark finds most agent optimizers erase their own gains over time.
-
278
Vint Cerf's Next Act: ID Cards for AI Agents
Internet pioneer Vint Cerf has joined Innovation Labs to push DNSid, a DNS-anchored identity standard for AI agents, through the IETF after retiring from Google.
-
277
Suno Hack Exposes 2 Million Scraped YouTube Songs
A hacker leaked Suno source code showing exactly how it scraped YouTube, Deezer, Genius, and a million hours of podcasts for AI training data.
-
276
Nvidia's H200 Chips Reach China - Congress Isn't Happy
Commerce official Jeffrey Kessler confirms H200 AI chip shipments to China have begun, but calls the volume "trivial" as lawmakers spar over export-control gaps.
-
275
How to Use an AI Browser Agent - A Beginner's Guide
A step-by-step guide to setting up your first AI browser agent, giving it a real task, and using it safely without handing over your passwords.
-
274
Hermes 4.3
Nous Research's 36B open-weight model matches Hermes 4 70B on most benchmarks, tops RefusalBench on alignment, and is the first production model trained entirely on the Solana-secured Psyche network.
-
273
Hassabis Calls for a US-Led Global AI Watchdog
Google DeepMind CEO Demis Hassabis wants a FINRA-style body testing frontier AI models before release, with power to slow the industry down if needed.
-
272
Kimi K2.7 Code Review: Open Weights Enter Copilot
Moonshot AI's Kimi K2.7-Code became the first open-weight model in GitHub Copilot's model picker, pairing genuine cost savings with a benchmark story that only Moonshot has verified.
-
271
OpenAI's Liability Shield Died - Illinois Passed Audits
OpenAI backed an Illinois bill shielding AI labs from mass-harm lawsuits, reversed course under public pressure, and watched the state sign a tougher audit law instead.
-
270
Hugging Face CEO Says Enterprises Are Done Renting AI
Clem Delangue says cost is pushing companies off frontier APIs and onto open models. A16z's own CIO survey shows enterprise dollars still moving the other way.
-
269
US Grants UAE License-Free Access to Nvidia AI Chips
The Commerce Department reclassified the UAE to Country Group A:5, giving NVIDIA, AMD, and Cerebras a clear path to supply AI chips and servers without per-shipment export licenses.
-
268
Mistral's Leanstral 1.5 Finds 5 Unknown Bugs Free
Mistral's open-source Leanstral 1.5 scanned 57 repos and found five previously unreported bugs, including a silent integer overflow in a Rust zigzag decoder.
-
267
Meta Cuts Muse Image in 3 Days After Creator Revolt
Meta's Muse Image let anyone generate AI images from public Instagram accounts by default. SAG-AFTRA and CAA pushed back. The feature lasted 72 hours.
-
266
Google Forces AI Label on Ads Ahead of EU Deadline
Google now requires advertisers to disclose AI-generated ad content across Search, YouTube, and Discover - three weeks before EU AI Act Article 50 enforcement begins.
-
265
Lyzr Raises $100M - Its Own AI Agent Did the Work
Lyzr closed a $100M Series B at a $500M valuation using its own AI agent SivaClaw to handle investor inquiries from 130+ funds - no founders needed on the road.
-
264
ByteDance Seedream 5.0 Pro - 2x Faster Than GPT-Image 2
ByteDance's Seedream 5.0 Pro reaches developer APIs with 2K output, 14-language text rendering, and layer-based editing at roughly one-fifth the cost of GPT-Image 2.
-
263
Ollama Banks $65M as Local AI Hits Enterprise Scale
Ollama raises $65M Series B led by Theory Ventures as 8.9 million monthly developers and 85% of Fortune 500 companies adopt local AI model deployment.
-
262
Grok 4.5 Ships - Token Efficiency Is the Real Edge
Grok 4.5 goes public with real benchmarks: behind Fable 5 on coding evals, but a 4.2x token efficiency gap that changes the cost math for high-volume pipelines.
-
261
Inside the 12-Day White House Gate on GPT-5.6 Sol
OpenAI's GPT-5.6 Sol goes public today after the first voluntary government hold on a frontier AI model - here's what the 12 days actually looked like.
-
260
ZML Ships a Free LLM Server That Runs on Any Chip
ZML's LLMD inference server runs Llama, Qwen, and Mistral on Nvidia, AMD, TPU, Apple Metal, and Intel Arc from a single binary - for free.
-
259
GPT-Live-1
OpenAI's full-duplex voice model that listens and speaks simultaneously, replacing Advanced Voice Mode in ChatGPT with three reasoning tiers backed by GPT-5.5.
-
258
DeepSeek Moves Into Chip Design to Beat Export Controls
DeepSeek is developing its own inference chip, confirmed by Reuters sources - the first semiconductor project in the lab's history, driven by US export controls that cut off China from Nvidia hardware and HBM memory.
-
257
UN AI Governance Forum: Bengio Warns of Catastrophic Risk
The UN's first all-nations AI governance dialogue opened in Geneva with Turing Award winner Yoshua Bengio warning that science cannot guarantee AI won't cause catastrophic harm.
We're indexing this podcast's transcripts for the first time — this can take a minute or two. We'll show results as soon as they're ready.
No matches for "" in this podcast's transcripts.
No topics indexed yet for this podcast.
Loading reviews...
ABOUT THIS SHOW
AI news, reviews, and analysis - discussed by two hosts who break down the latest in artificial intelligence, models, and agents. https://awesomeagents.ai/
HOSTED BY
Awesome Agents
CATEGORIES
Loading similar podcasts...