Data Skeptic cover art

All Episodes

Data Skeptic — 602 episodes

#
Title
1

Recommender Systems Today and Tomorrow

2

Recommender Systems Optimization Goals

3

Recommender Systems Origin Story

4

Social Choice for Fair Recommendations

5

News Recommendations

6

Give Users the Wheel

7

AutoLike

8

Student Spotlight: Aaron Payne, Data Analyst

9

The Future is Agentic in Recommender Systems

10

Book Ratings and Recommendations

11

Disentanglement and Interpretability in Recommender Systems

12

Collective Altruism in Recommender Systems

13

Niche vs Mainstream

14

Healthy Friction in Job Recommender Systems

15

Fairness in PCA-Based Recommenders

16

Video Recommendations in Industry

17

Eye Tracking in Recommender Systems

18

Cracking the Cold Start Problem

19

Designing Recommender Systems for Digital Humanities

20

DataRec Library for Reproducible in Recommend Systems

21

Shilling Attacks on Recommender Systems

22

Music Playlist Recommendations

23

Bypassing the Popularity Bias

24

Sustainable Recommender Systems for Tourism

25

Interpretable Real Estate Recommendations

26

Why Am I Seeing This?

27

Eco-aware GNN Recommenders

28

Networks and Recommender Systems

29

Network of Past Guests Collaborations

30

The Network Diversion Problem

31

Complex Dynamic in Networks

32

Github Network Analysis

33

Networks and Complexity

34

Graphs for Causal AI

35

Power Networks

36

Unveiling Graph Datasets

37

Network Manipulation

38

The Small World Hypothesis

39

Thinking in Networks

40

Fraud Networks

41

Criminal Networks

42

Graph Bugs

43

Organizational Network Analysis

44

Organizational Networks

45

Networks of the Mind

46

LLMs and Graphs Synergy

47

A Network of Networks

48

Auditing LLMs and Twitter

49

Fraud Detection with Graphs

50

Optimizing Supply Chains with GNN

51

The Mystery Behind Large Graphs

52

Customizing a Graph Solution

53

Graph Transformations

54

Networks for AB Testing

55

Lessons from eGamer Networks

56

Github Collaboration Network

57

Graphs and ML for Robotics

58

Graphs for HPC and LLMs

59

Graph Databases and AI

60

Network Analysis in Practice

61

Animal Intelligence Final Exam

62

Process Mining with LLMs

63

Open Animal Tracks

64

Bird Distribution Modeling with Satbird

65

Ant Encounters

66

Computing Toolbox

67

Biodiversity Monitoring

68

Hacking the Colony

69

Primate Poses

70

Generating 3D Animals with YouDream

71

Weird Communication

72

Reducing the Impact of Ship Noise on Marine Mammals

73

Analysis of Unstructured Data

74

iNaturalist

75

Learn to Code

76

Animal Computer Interaction

77

Ape Gestures

78

Evaluating AI Abilities

79

HMMs for Behavior

80

Bioinspired Engineering

81

Modelling Evolution

82

Behavioral Genetics

83

Signal in the Noise

84

Pose Tracking

85

Modeling Group Behavior

86

Advances in Data Loggers

87

What You Know About Intelligence is Wrong (fixed)

88

Animal Decision Making

89

Octopus Cognition

90

Optimal Foraging

91

Memory in Chess

92

OpenWorm

93

What the Antlion Knows

94

AI Roundtable

95

Uncontrollable AI Risks

96

I LLM and You Can Too

97

Q&A with Kyle

98

LLMs for Data Analysis

99

AI Platforms

100

Deploying LLMs

101

A Survey Assessing Github Copilot

102

Program Aided Language Models

103

Which Programming Language is ChatGPT Best At

104

GraphText

105

arXiv Publication Patterns

106

Do LLMs Make Ethical Choices

107

Emergent Deception in LLMs

108

Agents with Theory of Mind Play Hanabi

109

LLMs for Evil

110

The Defeat of the Winograd Schema Challenge

111

LLMs in Social Science

112

LLMs in Music Composition

113

Cuttlefish Model Tuning

114

Which Professions Are Threatened by LLMs

115

Why Prompting is Hard

116

Automated Peer Review

117

Prompt Refusal

118

A Long Way Till AGI

119

Brain Inspired AI

120

Computable AGI

121

AGI Can Be Safe

122

AI Fails on Theory of Mind Tasks

123

AI for Mathematics Education

124

Evaluating Jokes with LLMs

125

Why Machines Will Never Rule the World

126

A Psychopathological Approach to Safety in AGI

127

The NLP Community Metasurvey

128

Skeptical Survey Interpretation

129

The Gallup Poll

130

Inclusive Study Group Formation at Scale

131

The PhilPapers Survey

132

Non-Response Bias

133

Measuring Trust in Robots with Likert Scales

134

CAREER Prediction

135

The Panel Study of Income Dynamics

136

Survey Design Working Session

137

Bot Detection and Dyadic Surveys

138

Reproducible ESP Testing

139

A Survey of Data Science Methodologies

140

Opinion Dynamics Models

141

Casual Affective Triggers

142

Conversational Surveys

143

Do Results Generalize for Privacy and Security Surveys

144

4 out of 5 Data Scientists Agree

145

Crowdfunded Board Games

146

Russian Election Interference Effectiveness

147

Placement Laundering Fraud

148

Data Clean Rooms

149

Dark Patterns in Site Design

150

Internet Advertising Bureau Media Lab

151

Your Mouse Reveals Your Gender and Age

152

Measuring Web Search Behavior

153

StrategyQA and Big Bench

154

Ad Blockers Effect on News Consumption

155

Your Consent is Worth 75 Euros a Year

156

Automated Email Generation for Targeted Attacks

157

Tribal Marketing

158

Nano-targetted Facebook Ads

159

Debiasing GPT-3 Job Ads

160

ML Ops in Production

161

Ad Network Tomography

162

First Party Tracking Cookies

163

The Harms of Targeted Weight Loss Ads

164

Podcast Advertising

165

Fairness in e-Commerce Search

166

Fraudulent Amazon Reviewers

167

Ad Targeting in Amazon Smart Speakers

168

Adwords with Unknown Budgets

169

ML Ops Best Practices

170

Affiliate Marketing Rabbithole

171

Monetization of Youtube Conspiracy Theorists

172

User Perceptions of Problematic Ads

173

Political Digital Advertising Analysis

174

Fraud Detection in Crowdfunding Campaigns

175

Artificial Intelligence and Auction Design

176

Privacy Preference Signals

177

Neural Architecture Search for CTR Prediction

178

Algorithmic PPC Management

179

Data Skeptic: Ad Tech

180

The Reliability of Mobile Phone Data

181

Haywire Algorithms

182

School Reopening Analysis

183

Modern Data Stacks

184

Emoji as a Predictor

185

Polarizing Trends in the Gig Economy

186

Remote Learning in Applied Engineering

187

Remote Productivity

188

Does Remote Learning Work?

189

Covid-19 Impact on Bicycle Usage

190

Learning Digital Fabrication Remotely

191

Remote Software Development

192

Quantum K-Means

193

K-Means in Practice

194

Fair Hierarchical Clustering

195

Matrix Factorization For k-Means

196

Breathing K-Means

197

Power K-Means

198

Explainable K-Means

199

Customer Clustering

200

k-means Image Segmentation

201

Tracking Elephant Clusters

202

k-means clustering

203

Snowflake Essentials

204

Explainable Climate Science

205

Energy Forecasting Pipelines

206

Matrix Profiles in Stumpy

207

The Great Australian Prediction Project

208

Water Demand Forecasting

209

Open Telemetry

210

Fashion Predictions

211

Time Series Mini Episodes

212

Forecasting Motor Vehicle Collision

213

Deep Learning for Road Traffic Forecasting

214

Bike Share Demand Forecasting

215

Forecasting in Supply Chain

216

Black Friday

217

Aligning Time Series on Incomparable Spaces

218

Comparing Time Series with HCTSA

219

Change Point Detection Algorithms

220

Time Series for Good

221

Long Term Time Series Forecasting

222

Fast and Frugal Time Series Forecasting

223

Causal Inference in Educational Systems

224

Boosted Embeddings for Time Series

225

Change Point Detection in Continuous Integration Systems

226

Applying k-Nearest Neighbors to Time Series

227

Ultra Long Time Series

228

MiniRocket

229

ARiMA is not Sufficient

230

Comp Engine

231

Detecting Ransomware

232

GANs in Finance

233

Predicting Urban Land Use

234

Opportunities for Skillful Weather Prediction

235

Predicting Stock Prices

236

N-Beats

237

Translation Automation

238

Time Series at the Beach

239

Automatic Identification of Outlier Galaxy Images

240

Do We Need Deep Learning in Time Series

241

Detecting Drift

242

Darts Library for Time Series

243

Forecasting Principles and Practice

244

Prequisites for Time Series

245

Orders of Magnitude

246

They're Coming for Our Jobs

247

Pandemic Machine Learning Pitfalls

248

Flesch Kincaid Readability Tests

249

Fairness Aware Outlier Detection

250

Life May be Rare

251

Social Networks

252

The QAnon Conspiracy

253

Benchmarking Vision on Edge vs Cloud

254

Goodhart's Law in Reinforcement Learning

255

Video Anomaly Detection

256

Fault Tolerant Distributed Gradient Descent

257

Decentralized Information Gathering

258

Leaderless Consensus

259

Automatic Summarization

260

Gerrymandering

261

Even Cooperative Chess is Hard

262

Consecutive Votes in Paxos

263

Visual Illusions Deceiving Neural Networks

264

Earthquake Detection with Crowd-sourced Data

265

Byzantine Fault Tolerant Consensus

266

Alpha Fold

267

Arrow's Impossibility Theorem

268

Face Mask Sentiment Analysis

269

Counting Briberies in Elections

270

Sybil Attacks on Federated Learning

271

Differential Privacy at the US Census

272

Distributed Consensus

273

ACID Compliance

274

National Popular Vote Interstate Compact

275

Defending the p-value

276

Retraction Watch

277

Crowdsourced Expertise

278

The Spread of Misinformation Online

279

Consensus Voting

280

Voting Mechanisms

281

False Consensus

282

Fraud Detection in Real Time

283

Listener Survey Review

284

Human Computer Interaction and Online Privacy

285

Authorship Attribution of Lennon McCartney Songs

286

GANs Can Be Interpretable

287

Sentiment Preserving Fake Reviews

288

Interpretability Practitioners

289

Facial Recognition Auditing

290

Robust Fit to Nature

291

Black Boxes Are Not Required

292

Robustness to Unforeseen Adversarial Attacks

293

Estimating the Size of Language Acquisition

294

Interpretable AI in Healthcare

295

Understanding Neural Networks

296

Self-Explaining AI

297

Plastic Bag Bans

298

Self Driving Cars and Pedestrians

299

Computer Vision is Not Perfect

300

Uncertainty Representations

301

AlphaGo, COVID-19 Contact Tracing and New Data Set

302

Visualizing Uncertainty

303

Interpretability Tooling

304

Shapley Values

305

Anchors as Explanations

306

Mathematical Models of Ecological Systems

307

Adversarial Explanations

308

ObjectNet

309

Visualization and Interpretability

310

Interpretable One Shot Learning

311

Fooling Computer Vision

312

Algorithmic Fairness

313

Interpretability

314

NLP in 2019

315

The Limits of NLP

316

Jumpstart Your ML Project

317

Serverless NLP Model Training

318

Team Data Science Process

319

Ancient Text Restoration

320

ML Ops

321

Annotator Bias

322

NLP for Developers

323

Indigenous American Language Research

324

Talking to GPT-2

325

Reproducing Deep Learning Models

326

What BERT is Not

327

SpanBERT

328

BERT is Shallow

329

BERT is Magic

330

Applied Data Science in Industry

331

Building the howto100m Video Corpus

332

BERT

333

Onnx

334

Catastrophic Forgetting

335

Transfer Learning

336

Facebook Bargaining Bots Invented a Language

337

Under Resourced Languages

338

Named Entity Recognition

339

The Death of a Language

340

Neural Turing Machines

341

Data Infrastructure in the Cloud

342

NCAA Predictions on Spark

343

The Transformer

344

Mapping Dialects with Twitter Data

345

Sentiment Analysis

346

Attention Primer

347

Cross-lingual Short-text Matching

348

ELMo

349

BLEU

350

Simultaneous Translation at Baidu

351

Human vs Machine Transcription

352

seq2seq

353

Text Mining in R

354

Recurrent Relational Networks

355

Text World and Word Embedding Lower Bounds

356

word2vec

357

Authorship Attribution

358

Very Large Corpora and Zipf's Law

359

Semantic search at Github

360

Let's Talk About Natural Language Processing

361

Data Science Hiring Processes

362

Holiday Reading - Epicac

363

Drug Discovery with Machine Learning

364

Sign Language Recognition

365

Data Ethics

366

Escaping the Rabbit Hole

367

[MINI] Theorem Provers

368

Automated Fact Checking

369

[MINI] Single Source of Truth

370

Detecting Fast Radio Bursts with Deep Learning

371

Being Bayesian

372

Modeling Fake News

373

The Louvain Method for Community Detection

374

Cultural Cognition of Scientific Consensus

375

False Discovery Rates

376

Deep Fakes

377

Fake News Midterm

378

Quality Score

379

The Knowledge Illusion

380

Click Through Rates

381

Algorithmic Detection of Fake News

382

Ant Intelligence

383

Human Detection of Fake News

384

Spam Filtering with Naive Bayes

385

The Spread of Fake News

386

Fake News

387

Dev Ops for Data Science

388

First Order Logic

389

Blind Spots in Reinforcement Learning

390

Defending Against Adversarial Attacks

391

Transfer Learning

392

Medical Imaging Training Techniques

393

Kalman Filters

394

AI in Industry

395

AI in Games

396

Game Theory

397

The Experimental Design of Paranormal Claims

398

Winograd Schema Challenge

399

The Imitation Game

400

Eugene Goostman

401

The Theory of Formal Languages

402

The Loebner Prize

403

Chatbots

404

The Master Algorithm

405

The No Free Lunch Theorems

406

ML at Sloan Kettering Cancer Center

407

Optimal Decision Making with POMDPs

408

AI Decision-Making

409

[MINI] Reinforcement Learning

410

Evolutionary Computation

411

[MINI] Markov Decision Processes

412

Neuroscience Frontiers

413

Neuroimaging and Big Data

414

The Agent Model of Artificial Intelligence

415

Artificial Intelligence, a Podcast Approach

416

Holiday reading 2017

417

Complexity and Cryptography

418

Mercedes Benz Machine Learning Research

419

[MINI] Parallel Algorithms

420

Quantum Computing

421

Azure Databricks

422

[MINI] Exponential Time Algorithms

423

P vs NP

424

[MINI] Sudoku \in NP

425

The Computational Complexity of Machine Learning

426

[MINI] Turing Machines

427

The Complexity of Learning Neural Networks

428

[MINI] Big Oh Analysis

429

Data science tools and other announcements from Ignite

430

Generative AI for Content Creation

431

[MINI] One Shot Learning

432

Recommender Systems Live from FARCON 2017

433

[MINI] Long Short Term Memory

434

Zillow Zestimate

435

Cardiologist Level Arrhythmia Detection with CNNs

436

[MINI] Recurrent Neural Networks

437

Project Common Voice

438

[MINI] Bayesian Belief Networks

439

pix2code

440

[MINI] Conditional Independence

441

Estimating Sheep Pain with Facial Recognition

442

CosmosDB

443

[MINI] The Vanishing Gradient

444

Doctor AI

445

[MINI] Activation Functions

446

MS Build 2017

447

[MINI] Max-pooling

448

Unsupervised Depth Perception

449

[MINI] Convolutional Neural Networks

450

Multi-Agent Diverse Generative Adversarial Networks

451

[MINI] Generative Adversarial Networks

452

Opinion Polls for Presidential Elections

453

OpenHouse

454

[MINI] GPU CPU

455

[MINI] Backpropagation

456

Data Science at Patreon

457

[MINI] Feed Forward Neural Networks

458

Reinventing Sponsored Search Auctions

459

[MINI] The Perceptron

460

The Data Refuge Project

461

[MINI] Automated Feature Engineering

462

Big Data Tools and Trends

463

[MINI] Primer on Deep Learning

464

Data Provenance and Reproducibility with Pachyderm

465

[MINI] Logistic Regression on Audio Data

466

Studying Competition and Gender Through Chess

467

[MINI] Dropout

468

The Police Data and the Data Driven Justice Initiatives

469

The Library Problem

470

2016 Holiday Special

471

[MINI] Entropy

472

MS Connect Conference

473

Causal Impact

474

[MINI] The Bootstrap

475

[MINI] Gini Coefficients

476

Unstructured Data for Finance

477

[MINI] AdaBoost

478

Stealing Models from the Cloud

479

[MINI] Calculating Feature Importance

480

NYC Bike Share Rebalancing

481

[MINI] Random Forest

482

Election Predictions

483

[MINI] F1 Score

484

Urban Congestion

485

[MINI] Heteroskedasticity

486

Music21

487

[MINI] Paxos

488

Trusting Machine Learning Models with LIME

489

[MINI] ANOVA

490

Machine Learning on Images with Noisy Human-centric Labels

491

[MINI] Survival Analysis

492

Predictive Models on Random Data

493

[MINI] Receiver Operating Characteristic (ROC) Curve

494

Multiple Comparisons and Conversion Optimization

495

[MINI] Leakage

496

Predictive Policing

497

[MINI] The CAP Theorem

498

Detecting Terrorists with Facial Recognition?

499

[MINI] Goodhart's Law

500

Data Science at eHarmony

501

[MINI] Stationarity and Differencing

502

Feather

503

[MINI] Bargaining

504

deepjazz

505

[MINI] Auto-correlative functions and correlograms

506

Early Identification of Violent Criminal Gang Members

507

[MINI] Fractional Factorial Design

508

Machine Learning Done Wrong

509

Potholes

510

[MINI] The Elbow Method

511

Too Good to be True

512

[MINI] R-squared

513

Models of Mental Simulation

514

[MINI] Multiple Regression

515

Scientific Studies of People's Relationship to Music

516

[MINI] k-d trees

517

Auditing Algorithms

518

[MINI] The Bonferroni Correction

519

[MINI] Gradient Descent

520

Let's Kill the Word Cloud

521

2015 Holiday Special

522

Wikipedia Revision Scoring as a Service

523

[MINI] Term Frequency - Inverse Document Frequency

524

The Hunt for Vulcan

525

[MINI] The Accuracy Paradox

526

Neuroscience from a Data Scientist's Perspective

527

[MINI] Bias Variance Tradeoff

528

Big Data Doesn't Exist

529

[MINI] Covariance and Correlation

530

Bayesian A/B Testing

531

[MINI] The Central Limit Theorem

532

Accessible Technology

533

[MINI] Multi-armed Bandit Problems

534

[MINI] Structured and Unstructured Data

535

Measuring the Influence of Fashion Designers

536

[MINI] PageRank

537

Data Science at Work in LA County

538

[MINI] k-Nearest Neighbors

539

Crypto

540

[MINI] MapReduce

541

Genetically Engineered Food and Trends in Herbicide Usage

542

[MINI] The Curse of Dimensionality

543

Video Game Analytics

544

[MINI] Anscombe's Quartet

545

Proposing Annoyance Mining

546

Preserving History at Cyark

547

[MINI] A Critical Examination of a Study of Marriage by Political Affiliation

548

Detecting Cheating in Chess

549

[MINI] z-scores

550

Using Data to Help Those in Crisis

551

The Ghost in the MP3

552

Data Fest 2015

553

[MINI] Cornbread and Overdispersion

554

[MINI] Natural Language Processing

555

Computer-based Personality Judgments

556

[MINI] Markov Chain Monte Carlo

557

[MINI] Markov Chains

558

Oceanography and Data Science

559

[MINI] Ordinary Least Squares Regression

560

NYC Speed Camera Analysis with Tim Schmeier

561

[MINI] k-means clustering

562

Shadow Profiles on Social Networks

563

[MINI] The Chi-Squared Test

564

Mapping Reddit Topics with Randy Olson

565

[MINI] Partially Observable State Spaces

566

Easily Fooling Deep Neural Networks

567

[MINI] Data Provenance

568

Doubtful News, Geology, Investigating Paranormal Groups, and Thinking Scientifically with Sharon Hill

569

[MINI] Belief in Santa

570

Economic Modeling and Prediction, Charitable Giving, and a Follow Up with Peter Backus

571

[MINI] The Battle of the Sexes

572

The Science of Online Data at Plenty of Fish with Thomas Levi

573

[MINI] The Girlfriend Equation

574

The Secret and the Global Consciousness Project with Alex Boklin

575

[MINI] Monkeys on Typewriters

576

Mining the Social Web with Matthew Russell

577

[MINI] Is the Internet Secure?

578

Practicing and Communicating Data Science with Jeff Stanton

579

[MINI] The T-Test

580

Data Myths with Karl Mamer

581

Contest Announcement

582

[MINI] Selection Bias

583

[MINI] Confidence Intervals

584

[MINI] Value of Information

585

Game Science Dice with Louis Zocchi

586

Data Science at ZestFinance with Marick Sinay

587

[MINI] Decision Tree Learning

588

Jackson Pollock Authentication Analysis with Kate Jones-Smith

589

[MINI] Noise!!

590

Guerilla Skepticism on Wikipedia with Susan Gerbic

591

[MINI] Ant Colony Optimization

592

Data in Healthcare IT with Shahid Shah

593

[MINI] Cross Validation

594

Streetlight Outage and Crime Rate Analysis with Zach Seeskin

595

[MINI] Experimental Design

596

The Right (big data) Tool for the Job with Jay Shankar

597

[MINI] Bayesian Updating

598

Personalized Medicine with Niki Athanasiadou

599

[MINI] p-values

600

Advertising Attribution with Nathan Janos

601

[MINI] type i / type ii errors

602

Introduction