Fast, no‑friction execution across every stage of code‑data collection and annotation — from pre‑training to post‑training.
30 engineers, annotators, and domain experts. Rigorous hiring pipeline — up to 100 structured interviews per week.
Cross‑validation by practicing industry experts. Multi‑stage QC gates and per‑rater calibration.
Compliance with data security and confidentiality standards. Full PII redaction across enterprise data.
Not just a data supplier — we own every stage from sourcing to evaluation. Plug us in at any step, or hand off the whole chain.
Non‑public repos, real enterprise content, licensed archives — never seen in public sets.
Prompt → response pairs, multi‑turn dialogues, code‑task authoring at scale.
Human‑in‑the‑loop labeling, rationales, pairwise ranking across 40 criteria.
SWE‑Bench, Multi‑SWE, Harbor, RAG eval, Dockerized reproducible environments.
Agent‑trajectory scoring, plan/thought eval, safety probes, regression testing.
Non‑public code repositories, SWE‑Bench benchmarks, alignment sets, and enterprise data.
Production code from real companies — never indexed, never crawled.
4,200+ proprietary repositories from 1,900+ companies, 580M+ lines of code — none of it ever appeared in public training sets (GitHub, GitLab, HuggingFace). Production‑grade repositories from real companies — primarily sourced from a network of outsourcing agencies and startups whose products were discontinued or acquired.
By repository: PHP 19%, TypeScript 16%, JavaScript 15%, Python 11%, Swift 6%, Java 5%, Kotlin 5%, other 23%.
Composition: 67% discontinued / 33% active or maintained. Full legal rights to license every repository.
Internal enterprise data sourced from real companies — active, acquired, or wound down — each with certified consent to license.
Not just data delivery — our engineers integrate at every stage of your pipeline. Standard formats, your evaluation frameworks, schema design, ingestion adapters, continuous delivery on your cadence — data arrives ready to train or benchmark.
JSONL for SFT and DPO, Parquet for pre‑training corpora, HuggingFace Datasets for publishing. Conversation data in ShareGPT, Alpaca, or OpenAI chat schemas — drop‑in for your training loop, no conversion on your side.
Dockerized SWE‑Bench and Multi‑SWE‑Bench harnesses (Harbor‑compatible), RAG eval — patch‑apply, build, and test runs work identically on our machines and yours.
Practicing developers, ML engineers, and domain specialists as an extension of your team — scoping taxonomies, defining quality criteria, resolving edge cases as they surface.
Expanded and improved version of the agent quality standard
16.04.25One consistent quality standard, no matter what you code in
24.12.24Cutting errors by 40% and costs by 60%
Who builds the data, how a first project runs, and the two things every team asks about the non‑public repository collection.
We came to AI data from software development, not the other way round. The company wrote production code first and started its ML and AI work in 2017. Today we work with frontier labs on coding data and model evaluation.
No. Every coding task is written and reviewed by our own middle and senior engineers. We match the work to background, so an ML task goes to an ML engineer and an Android task to someone who ships Android.
Both. You can license data and benchmark tasks we already have, or commission a set built around a specific language, skill, format or difficulty level. The cheaper order is usually to test an existing sample first: once we can see where your model actually struggles, the custom work goes where it pays off.
The rubric and the baseline are agreed before production starts, never after. Depending on the task, checks can include automated tests, model-based review, blind comparison, review by a developer, and metrics from the model downstream. We normally start with a POC, so you can weigh quality, cost and the effect on the model before committing to a bigger run.
Whole codebases from real companies, with the project structure and commit history intact. None of it was ever published as open source. The languages run to JavaScript, TypeScript, PHP, Python, Objective‑C and Java among others, and the products behind them come from banking, retail, healthcare, logistics, education and IT services.
Yes. Every repository comes to us under an agreement that lets us license the data, and the rights holders are paid royalties on it. We can show the supporting paperwork under NDA. Ownership of the original code, our right to license it, and any exclusivity you want are three separate things, and the contract spells out each one.
Curated subset of non‑public repositories and benchmark tasks for hands‑on quality validation. Schedule a technical deep dive with our engineering team.
Email: hi@fermatix.ai
AVENIDAS INTELIGENTES, LDA
Lg Alberto Sampaio, 3 A, Sala 10
Linda a Velha, 2795‑007
Portugal