LLM, AI Data Analyst, Exploratory Data Analysis (EDA), AI-assisted analytics, tabular data analysis, LLM evaluation, data cleaning and AI agents—how well do they actually work together?
- GitHub / Benchmark: https://github.com/deepsense-ai/eda-b…
- Read more: https://deepsense.ai/blog/synthetic-d…
In this Tech Experts Webinar, Szymon Betlewski, ML Engineer, presents an EDA benchmark designed to evaluate how well today’s AI agents perform exploratory data analysis on realistic tabular datasets.
Unlike clean benchmark datasets, the tasks simulate common real-world data issues, including missing values, inconsistent formats, duplicate records, hidden constraints, and noisy data.
The webinar covers:
- how the EDA benchmark was designed
- synthetic datasets for reproducible evaluation
- benchmark tasks across multiple business domains
- evaluation methodology and scoring
- strengths and weaknesses of current LLMs
- reliability-adjusted scoring
- leaderboard comparison across leading models
The session also discusses where AI agents already perform well—and where careful reasoning and consistent decision-making still require human expertise.
Timeline
00:00 Intro & Why EDA is difficult
03:18 Building the EDA benchmark
06:01 Benchmark methodology
09:13 LLM leaderboard and results
11:17 Key findings and conclusions
Speaker
Szymon Betlewski
ML Engineer






