Home / Founders / Directory / Data & labeling
Data & Labeling Founders
Profiles of the 19 founders building tools that prepare, label and generate training data, and the retrieval systems behind AI models, listed A–Z by last name.
19 founders
Bakouk is building Sifflet so data engineers and the people who use their dashboards see the same alert, lineage, and context when a number breaks.
Coe is trying to make sensitive production data safe for developers to use, so engineering teams can build and test software without waiting on access to the real thing.
Eskildsen built turbopuffer after doing the math on what vector search cost, and set out to make search over every byte affordable.
Handy turned an analytics-consultancy workflow into dbt, the default way many teams transform warehouse data as code.
Hansen is pushing multimodal AI teams to treat data curation and annotation as core infrastructure, not a spreadsheet side job.
Kirwan turned Uber-scale experimentation pain into Bigeye, a data observability and AI trust platform for enterprises that cannot afford silent pipeline failures.
Liberty built Pinecone so long-term memory for AI is infrastructure product teams can buy, not rebuild as research.
van Luijt built Weaviate so search over meaning stays open infrastructure teams can run themselves.
Masschelein is building Soda to catch bad data where it is produced, with checks and data contracts written by the engineers who own the pipelines.
Petrosyan is building SuperAnnotate around the expert human judgment that AI labs and enterprises need to label data and grade model answers.
Ramdas is trying to shrink the years between a promising semiconductor material and something a fab will put on a line.
Raymond is focused on the unglamorous bottleneck before retrieval: turning messy enterprise files into model-ready data.
Segal wants data teams to run ingestion, reverse ETL, observability and cataloging in one system, so people asking AI tools questions can trust the data underneath.
Sharma built labeling software so computer-vision teams can train models without reinventing annotation infrastructure every time.
Shmukler is bringing product-and-growth instincts from LinkedIn, Wealthfront, and Instacart to Anomalo's bet that AI can catch unknown data bugs before dashboards lie.
Tricot wants data integration pipelines to become a commodity teams can extend, instead of a one-off project for every SaaS source.
Vasnetsov built Qdrant because existing vector libraries were not enough for production similarity search at scale.
Wang built Scale around the idea that frontier AI is gated by high-quality data and evaluation, not just bigger models.
Xie built Zilliz and open-sourced Milvus to make vector search a first-class database category for AI applications.
No founders match that search.
