EPISODE · Sep 19, 2026 · 15 MIN
I Tested Gemini 3 Against GPT-4. The Results Shocked Me.
from Unboxed · host James Caldwell
Google just dropped Gemini 3 DeepThink, and the AI world is scrambling to figure out what just happened. While everyone was watching OpenAI's latest updates, Google quietly released something that's making GPT-4 look like last year's model. The numbers are pretty wild. Gemini 3 DeepThink scored 94.2% on MMLU benchmarks compared to GPT-4's 86.4% and Claude 3.5 Sonnet's 88.7%. That's not a small jump. This isn't just Google catching up anymore. But here's what's really interesting: DeepThink uses up to 10x more compute per query than standard Gemini 3. Response times are significantly slower, but the reasoning capabilities show a 67% improvement on mathematical tasks. Google's basically trading speed for accuracy, which tells us something important about where AI is heading. James spent the weekend testing DeepThink against GPT-4 on complex reasoning problems, and the results surprised him. This isn't just benchmark optimization. The model approaches multi-step problems differently, and it shows. In This Episode: > How DeepThink's architecture differs from standard language models > Real-world testing results on coding, math, and logical reasoning tasks > What this means for developers currently building on OpenAI's API > Why Google released this as a limited preview instead of full rollout Timestamps: 00:00 Introduction to Gemini 3 DeepThink 02:15 Benchmark results breakdown 04:30 Head-to-head testing methodology 06:45 Complex reasoning task comparisons 08:20 What this means for AI development 10:30 Implications for current AI users Google's making a serious play for the reasoning crown. If you're building anything that requires complex problem-solving, this episode breaks down what you need to know about the new AI landscape. Follow Unboxed for daily AI updates that actually matter. New episodes drop multiple times daily because this space moves fast. Learn more about your ad choices. Visit megaphone.fm/adchoices
Embed this episode
Ready to play
I Tested Gemini 3 Against GPT-4. The Results Shocked Me.
No transcript for this episode yet
Similar Episodes
No similar episodes found.