Inference Scaling for Long-Context RAG episode artwork

EPISODE · Oct 20, 2024 · 12 MIN

Inference Scaling for Long-Context RAG

from LlamaCast · host Shahriar Shariati

🗓 Inference Scaling for Long-Context Retrieval Augmented GenerationThis research paper explores the effectiveness of inference scaling for retrieval augmented generation (RAG), a technique that enhances large language models (LLMs) by incorporating external knowledge. The authors introduce two strategies, demonstration-based RAG (DRAG) and iterative demonstration-based RAG (IterDRAG), for effectively scaling inference computation. They demonstrate that increasing inference computation, when optimally allocated, leads to nearly linear gains in RAG performance. Furthermore, they develop a computation allocation model to predict the optimal test-time compute allocation for various tasks and scenarios, showcasing its effectiveness in achieving performance gains and aligning with experimental results.📎 Link to paper

Episode metadata supplied by the publisher feed · Published Oct 20, 2024

Embed this episode

Ready to play

Inference Scaling for Long-Context RAG

0:00 12:18

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of LlamaCast?

This episode is 12 minutes long.

When was this LlamaCast episode published?

This episode was published on October 20, 2024.

Can I download this LlamaCast episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!