About deepsense.ai
deepsense.ai is a 120-person AI/ML consultancy. We’re an OpenAI Advanced Partner and an Anthropic Service Partner, with a direct line to both labs’ solution engineers. We also build production work on ElevenLabs, so you’re getting early, hands-on access to new models and features before they hit the mainstream. For over a decade we’ve delivered applied AI projects for companies like J&J, Sky, John Deere, and GLS — spanning LLM applications, agents, MLOps, and data science. We’re a people-first organization that believes in deep technical craft and real ownership. Our engineers don’t just advise; they build things that run in production.
The opportunity
At deepsense.ai, we design challenging datasets, problems, and evaluation frameworks for state-of-the-art LLMs and VLMs. We want to understand what advanced AI models can actually do, where they fail, and what to test next.
We’re looking for an AI Benchmark & Datasets Engineer to help us design and build new benchmarks for evaluating frontier AI models. Our current focus is Exploratory Data Analysis (EDA) and broader Data Science reasoning — but part of the role is exploring what should come next.
Datasets may be fully synthetic or based on real-world data. The real challenge isn’t preparing data — it’s designing problems that require genuine reasoning and reliably separate stronger models from weaker ones.
What you’ll do
- Design original benchmark problems (EDA, Data Science, and related analytical domains) and build the datasets behind them, synthetic or real.
- Define ground truth, expected outputs, and evaluation criteria; write clear, reproducible task specifications.
- Test problems against current models, analyze failure modes and shortcuts, and calibrate difficulty.
- Review and evaluate benchmark problems created by other contributors.
- Investigate new classes of problems and model capabilities worth testing — and document everything in English.
What we’re looking for
- Solid understanding of Data Science and Exploratory Data Analysis, with hands-on experience working with data.
- Working knowledge of Python (scripting, basic analysis) and the Data Science ecosystem.
- Working knowledge of Git (clone, commit, basic branching).
- Strong analytical thinking — the ability to independently explore unfamiliar datasets and formulate non-obvious, interesting questions.
- Creativity in designing varied, non-standard test scenarios, plus attention to detail and methodological rigor.
- Good written English, at least B2.
You don’t need to already be an LLM evaluation expert — what matters most is the ability to take an unfamiliar problem, understand it deeply, and turn it into a rigorous, evaluable challenge.
Nice to have
- Experience with LLMs/VLMs, benchmark design, or synthetic data generation.
- Background in ML/Data Science research, experimental design, or statistical analysis.
- Kaggle or analytical competitions, academic research, or open-source projects.
- Experience creating educational or assessment materials.
What working together looks like
- You pick up benchmark tasks — designing a new problem from scratch, or reviewing and evaluating one built by someone else.
- There’s no fixed number of tasks required — you decide how much you take on and at what pace.
- Once you take on a task, you’ll have a few days to complete it and submit it for review.
- Everything is 100% remote, on a contract of mandate (umowa zlecenie).
Recruitment process
- Application — send your CV and, if you have them, links to relevant projects, repos, or publications.
- Practical task — a take-home assignment, with 3–5 days to complete and submit it.
- 30-minute call — we discuss your solution and how you approached it.
Interested?
If you enjoy working with data, solving hard problems, and probing the limits of modern AI systems, we’d like to hear from you. You don’t need to come from an AI research lab — we’re looking for people who can understand complex problems, invent good ones, and turn them into rigorous evaluations.
