Gen AI Series: Legalities of scraping data to build LLMs, a conversation w/ Adam Shevell, partner at Wilson Sonsini episode artwork

EPISODE · Nov 1, 2024 · 45 MIN

Gen AI Series: Legalities of scraping data to build LLMs, a conversation w/ Adam Shevell, partner at Wilson Sonsini

from From Startup to Exit · host TiE Seattle

Send us Fan MailMany of the large language model providers have built their LLM models by scraping data from websites and open source sites. Some have licensed the content like Open AI has done with Reddit. Nonetheless, many of these LLM providers have been sued for illegally copying data for which they have no permission. In this podcast, Adam shares his thoughts on the Fair Use doctrine and various legal opinions on the legality of scraping data to build LLM models. Here's also an interesting article that Adam has written for developing, extending and using Generative AI models.Adam Shevell is a partner in Wilson Sonsini’s San Francisco office, where he co-leads the firm’s technology transactions practice in the city. Adam advises technology companies and their investors at all stages of company development, from pioneering start-ups to leading global enterprises, angel investors, venture capital firms, and other institutions in the start-up ecosystem. Adam represents leading Silicon Valley companies on complex and strategic transactions involving cutting-edge innovations and the launch of new products. By understanding his clients’ products, markets, and business priorities, and by building deep and lasting relationships, Adam provides creative and pragmatic advice focused on delivering effective results.Adam also works closely with Canadian start-ups on their U.S. expansion, fundraising, strategic partnerships, and exit transactions, and with Canadian venture funds investing in U.S.-based companies.Brought to you by TiE SeattleHosts: Shirish Nadkarni and Gowri ShankarProducers: Minee Verma and Eesha JainYouTube Channel: https://www.youtube.com/@fromstartuptoexitpodcast

Episode metadata supplied by the publisher feed · Published Nov 1, 2024

Embed this episode

Send us Fan Mail Many of the large language model providers have built their LLM models by scraping data from websites and open source sites. Some have licensed the content like Open AI has done with Reddit. Nonetheless, many of these LLM providers have been sued for illegally copying data for which they have no permission. In this podcast, Adam shares his thoughts on the Fair Use doctrine and various legal opinions on the legality of scraping data to build LLM models. Here's also an interest...

Distinct summary based on available episode metadata or transcript content.

NOW PLAYING

Gen AI Series: Legalities of scraping data to build LLMs, a conversation w/ Adam Shevell, partner at Wilson Sonsini

0:00 45:12

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

No similar podcasts found.

Frequently Asked Questions

How long is this episode of From Startup to Exit?

This episode is 45 minutes long.

When was this From Startup to Exit episode published?

This episode was published on November 1, 2024.

Is there a transcript available for this episode?

Yes, a full transcript is available for this episode. You can read the complete transcript on the episode page.

Can I download this From Startup to Exit episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!