On the (Mis)Use of Machine Learning with Panel Data episode artwork

EPISODE · Jul 25, 2025 · 17 MIN

On the (Mis)Use of Machine Learning with Panel Data

from Marketing^AI · host Enoch H. Kang

This academic paper investigates the critical issue of data leakage in applying machine learning (ML) to panel data, which combines cross-sectional and time-series observations. The authors explain that standard ML practices, when unsuited for panel data's inherent structure, can lead to temporal leakage (future information affecting past predictions) and cross-sectional leakage (information sharing across training and testing units). This leakage results in inflated model performance and misleading policy recommendations, as empirical applications, particularly for income prediction in U.S. counties, vividly demonstrate. To counter this, the paper offers practical guidelines for practitioners, emphasizing the importance of clearly defining research goals—whether for cross-sectional prediction or sequential forecasting—and implementing appropriate data splitting and cross-validation strategies to ensure robust and realistic ML model evaluation.

Episode metadata supplied by the publisher feed · Published Jul 25, 2025

Embed this episode

Ready to play

On the (Mis)Use of Machine Learning with Panel Data

0:00 17:37

No transcript for this episode yet

We transcribe on demand. Request one and we'll notify you when it's ready — usually under 10 minutes.

No similar episodes found.

Frequently Asked Questions

How long is this episode of Marketing^AI?

This episode is 17 minutes long.

When was this Marketing^AI episode published?

This episode was published on July 25, 2025.

Can I download this Marketing^AI episode?

Yes. Use the download control on the episode player to save the publisher-provided media file.
URL copied to clipboard!