← Home FAQ

Frequently asked questions

The questions AI labs ask us before the first pilot: where the data comes from, what you are allowed to do with it, and how a custom project gets scoped.

About Fermatix

Why should we choose Fermatix?

We came to AI data from software development, not the other way round. The company wrote production code first and started its ML and AI work in 2017. Today we work with frontier labs on coding data and model evaluation.

Who builds the datasets?

Middle and senior developers on our own staff, around 30 engineers, AI trainers and expert annotators, covering 21 programming languages between them. Work is assigned by stack rather than by who happens to be free, and the team spans time zones, so production runs around the clock.

Do you use crowdsourcing?

No. Every coding task is written and reviewed by our own middle and senior engineers. We match the work to background, so an ML task goes to an ML engineer and an Android task to someone who ships Android.

How do you scale a project up?

Scale comes out of hiring rather than a crowd platform. The pipeline runs up to 100 structured interviews a week, which is what takes a project from a first sample to 12,000+ datapoints a month.

What happens to the work before it reaches us?

It goes through multi‑level quality control and cross‑validation, and the people doing that review are practicing engineers. What else gets checked depends on the rubric agreed at the start of the project.

Licensed repository data

What is included in your private repository dataset?

Whole codebases from real companies, with the project structure and commit history intact. None of it was ever published as open source. The languages run to JavaScript, TypeScript, PHP, Python, Objective‑C and Java among others, and the products behind them come from banking, retail, healthcare, logistics, education and IT services.

How do you protect sensitive information?

Every repository is reviewed twice before delivery, once automatically and once by a person. We strip PII and secrets, drop files that have no business being in a training set, and measure how much third‑party open‑source code the project carries. How far that goes depends on what you plan to do with the data, so the cleaning rules get agreed with you up front.

Does the dataset contain AI‑generated code?

No. What is in there is production code written by human teams. We confirm the origin with the repository owners, and projects containing AI‑generated code do not go into the set.

Do you have the right to license the repositories?

Yes. Every repository comes to us under an agreement that lets us license the data, and the rights holders are paid royalties on it. We can show the supporting paperwork under NDA. Ownership of the original code, our right to license it, and any exclusivity you want are three separate things, and the contract spells out each one.

Is the data exclusive?

We source the repositories ourselves instead of buying them from brokers, and we do not resell the collection through other data companies. That said, a repository can end up licensed to more than one lab. If you need that ruled out, exclusivity is negotiable for a defined dataset.

What licensing options are available?

Fixed‑term or perpetual, both are on the table. Price moves with the size of the set, how narrow the selection criteria are, how much preparation the data needs, the length of the license and whether you want exclusivity. A large package costs less per unit than a small handpicked subset. We put final numbers on it once the dataset and the usage rights are settled.

Custom datasets and benchmarks

Can you tailor the dataset to our needs?

Yes. Repositories can be filtered by language, industry, project size, product type and other metadata, and they can be turned into training tasks, benchmarks or evaluation environments aimed at whatever your model is weakest at. Most projects start with a sample: we run it against your model, read the failures, and let that decide what gets built next.

What coding and agent benchmarks can you build?

Repository‑level engineering, black‑box implementation, multi‑tool workflows, iterative optimisation and end‑to‑end research. Recent formats such as DeepSWE, FrontierSWE, ProgramBench, APEX‑Agents, ALE, Frontier‑Eng and ResearchClawBench serve as references. The tasks can be built from private repositories, from open source, or from your own code and data, and we set the difficulty against baseline models and a target pass rate agreed with you.

Do you offer ready‑made data or custom projects?

Both. You can license data and benchmark tasks we already have, or commission a set built around a specific language, skill, format or difficulty level. The cheaper order is usually to test an existing sample first: once we can see where your model actually struggles, the custom work goes where it pays off.

How do you evaluate quality?

The rubric and the baseline are agreed before production starts, never after. Depending on the task, checks can include automated tests, model-based review, blind comparison, review by a developer, and metrics from the model downstream. We normally start with a POC, so you can weigh quality, cost and the effect on the model before committing to a bigger run.

Still have a question?

Tell us what you are training or evaluating, and we will point you at the closest sample.