Evaluating the Effectiveness of Large Language Models: Challenges and Insights // Aniket Singh // #248 episode artwork

EPISODE · Jul 16, 2024 · 35 MIN

Evaluating the Effectiveness of Large Language Models: Challenges and Insights // Aniket Singh // #248

from MLOps.community · host Demetrios

Aniket Kumar Singh is a Vision Systems Engineer at Ultium Cells, skilled in Machine Learning and Deep Learning. I'm also engaged in AI research, focusing on Large Language Models (LLMs).Evaluating the Effectiveness of Large Language Models: Challenges and Insights // MLOps Podcast #248 with Aniket Kumar Singh, CTO @ MyEvaluationPal | ML Engineer @ Ultium Cells.// AbstractDive into the world of Large Language Models (LLMs) like GPT-4. Why is it crucial to evaluate these models, how we measure their performance, and the common hurdles we face? Drawing from Aniket's research, he shares insights on the importance of prompt engineering and model selection. Aniket also discusses real-world applications in healthcare, economics, and education, and highlights future directions for improving LLMs.// BioAniket is a Vision Systems Engineer at Ultium Cells, skilled in Machine Learning and Deep Learning. I'm also engaged in AI research, focusing on Large Language Models (LLMs).// MLOps Jobs board jobs.mlops.community// MLOps Swag/Merchhttps://mlops-community.myshopify.com/// Related LinksWebsite: www.aniketsingh.meAniket's AI Research for Good blog that I plan to utilize to share any new research that would focus on the good: www.airesearchforgood.orgAniket's papers: https://scholar.google.com/citations?user=XHxdWUMAAAAJ&hl=en --------------- ✌️Connect With Us ✌️ -------------Join our Slack community: https://go.mlops.community/slackFollow us on Twitter: @mlopscommunitySign up for the next meetup: https://go.mlops.community/registerCatch all episodes, blogs, newsletters, and more: https://mlops.community/Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/Connect with Aniket on LinkedIn: https://www.linkedin.com/in/singh-k-aniket/Timestamps:[00:00] Aniket's preferred coffee[00:14] Takeaways[01:29] Aniket's job and hobby[03:06] Evaluating LLMs: Systems-Level Perspective[05:55] Rule-based system[08:32] Evaluation Focus: Model Capabilities[13:04] LLM Confidence[13:56] Problems with LLM Ratings[17:17] Understanding AI Confidence Trends[18:28] Aniket's papers[20:40] Testing AI Awareness[24:36] Agent Architectures Overview[27:05] Leveraging LLMs for tasks[29:53] Closed systems in Decision-Making[31:28] Navigating model Agnosticism[33:47] Robust Pipeline vs Robust Prompt[34:40] Wrap up

Aniket Kumar Singh is a Vision Systems Engineer at Ultium Cells, skilled in Machine Learning and Deep Learning. I'm also engaged in AI research, focusing on Large Language Models (LLMs).Evaluating the Effectiveness of Large Language Models: Challenges and Insights // MLOps Podcast #248 with Aniket Kumar Singh, CTO @ MyEvaluationPal | ML Engineer @ Ultium Cells.// AbstractDive into the world of Large Language Models (LLMs) like GPT-4. Why is it crucial to evaluate these models, how we measure their performance, and the common hurdles we face? Drawing from Aniket's research, he shares insights on the importance of prompt engineering and model selection. Aniket also discusses real-world applications in healthcare, economics, and education, and highlights future directions for improving LLMs.// BioAniket is a Vision Systems Engineer at Ultium Cells, skilled in Machine Learning and Deep Learning. I'm also engaged in AI research, focusing on Large Language Models (LLMs).// MLOps Jobs board jobs.mlops.community// MLOps Swag/Merchhttps://mlops-community.myshopify.com/// Related LinksWebsite: www.aniketsingh.meAniket's AI Research for Good blog that I plan to utilize to share any new research that would focus on the good: www.airesearchforgood.orgAniket's papers: https://scholar.google.com/citations?user=XHxdWUMAAAAJ&hl=en --------------- ✌️Connect With Us ✌️ -------------Join our Slack community: https://go.mlops.community/slackFollow us on Twitter: @mlopscommunitySign up for the next meetup: https://go.mlops.community/registerCatch all episodes, blogs, newsletters, and more: https://mlops.community/Connect with Demetrios on LinkedIn: https://www.linkedin.com/in/dpbrinkm/Connect with Aniket on LinkedIn: https://www.linkedin.com/in/singh-k-aniket/Timestamps:[00:00] Aniket's preferred coffee[00:14] Takeaways[01:29] Aniket's job and hobby[03:06] Evaluating LLMs: Systems-Level Perspective[05:55] Rule-based system[08:32] Evaluation Focus: Model Capabilities[13:04] LLM Confidence[13:56] Problems with LLM Ratings[17:17] Understanding AI Confidence Trends[18:28] Aniket's papers[20:40] Testing AI Awareness[24:36] Agent Architectures Overview[27:05] Leveraging LLMs for tasks[29:53] Closed systems in Decision-Making[31:28] Navigating model Agnosticism[33:47] Robust Pipeline vs Robust Prompt[34:40] Wrap up

NOW PLAYING

Evaluating the Effectiveness of Large Language Models: Challenges and Insights // Aniket Singh // #248

0:00 35:40

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

She’s a Hazard to Herself She’s a Hazard Hi there, I’m Mallory, and I’d like to invite you into our world with “She’s a Hazard to Herself!” Join us as we navigate life with Multiple Sclerosis from the seat of my power wheelchair. Discover stories of resilience, family, and the community we’ve built around chronic illness. Whether you’re impacted by MS or want to learn from our journey, there’s something here for you. So why wait? Subscribe to “She’s a Hazard to Herself” on your favorite podcast app and be part of our journey today. Let’s lift each other up, one episode at a time! Tips, News and Stories for Older Adults Esther C Kane CAPS, C.D.S. "Tips, News, and Stories for Older Adults" delivers weekly insights tailored for seniors. We bring you summaries of curated news, practical advice, and inspiring stories that matter to the 55+ community. From health and finance to technology and lifestyle, our content keeps you informed and engaged. Sourced from trusted outlets, each episode offers valuable information for navigating your golden years. Join us as we explore aging with positivity, wisdom, and engaging stories. Your perfect companion for staying active, learning, and embracing life's later chapters. Prayer Time Heir Waves Prayer Time A podcast especially for our Prayer Time community NEWMORROW SESSIONS - A PodCast Series on the Future of Hospitality Mario C. Bauer, Florian Schneider, Axel Weber & Dr. Tillman Bardt The Newmorrow PodCast is more than a podcast — it's a platform for open dialog on the future of our business, a platform for those building what doesn’t exist yet. Here, we share and embrace our passion for the hospitality industry, but we won’t romanticize the journey. We ask the tough questions, confront uncomfortable truths, and prepare for a future that resists easy answers. We believe that the tougher and wilder times become, the more openly, honestly and humanely people need to talk to each other and act together. We believe, openness, togetherness, and truthfulness should also be cornerstones of a professional community to develop our utopian idea of „open source“. This is a space where visionaries don’t just imagine the future — they wrestle with the paradoxes that shape it: success vs. happiness, data vs. instinct, stability vs. reinvention. Join leaders, entrepreneurs, and thinkers as they share not what made them — but what’s actively shaping them, now and next. So tune in

Frequently Asked Questions

How long is this episode of MLOps.community?

This episode is 35 minutes long.

When was this MLOps.community episode published?

This episode was published on July 16, 2024.

What is this episode about?

Aniket Kumar Singh is a Vision Systems Engineer at Ultium Cells, skilled in Machine Learning and Deep Learning. I'm also engaged in AI research, focusing on Large Language Models (LLMs).Evaluating the Effectiveness of Large Language Models:...

Can I download this MLOps.community episode?

Yes, you can download this episode by clicking the download button on the episode player, or subscribe to the podcast in your preferred podcast app for automatic downloads.
URL copied to clipboard!