Multi-stage data analysis in genetics and quantitative biology. Agents get messy, staged datasets and a minimal prompt. They must handle QC, selection bias, confounding and model choice to recover a verifiable target estimand. The problems span statistical, population, quantitative and functional genetics, plus clinical, cancer, proteomics, spatial transcriptomics, epigenomics and forensic genetics.
This sample task is provided for illustration only. The scores may not reflect a model's overall performance on this benchmark.
You are given allele-frequency time series data from two haploid loci sampled over multiple generations.
One locus is under stronger positive selection than the other. Estimate the selection coefficient s for the more strongly selected locus, where s > 0 means the derived allele is favored.
Assume instrument-driven sequencing error is ~1%. The seq_error column is the average of the two directional allele-miscall rates for that locus and sample.
The selected_locus value must be "A" or "B".
These data came from a real experiment; you will be graded not just on numerical correctness but the quality of analytical reasoning you exhibit; do not attempt to take any shortcuts.
Return your final answer as exactly one JSON object.
Do not wrap the JSON in markdown.
Do not add prose before or after the JSON.
Do not omit any keys shown in the example.
Return the JSON object in your final answer:
{
"answer": {
"selected_locus": "<string>",
"s": <float>
},
"reasoning": "<description of method and QC>"
}
The data files for this problem are:
/workspace/data_files/locus_A_timeseries.tsv.gz
/workspace/data_files/locus_B_timeseries.tsv.gz
/workspace/data_files/variant_info.tsv.gz
/workspace/data_files/Ne_schedule.tsv.gz
Also save exactly that JSON object to /workspace/eval_answer.json. Only that file is graded.
No rubric data available for this model.
APEX NEWSLETTER
The latest on frontier AI performance, straight to your inbox.
New benchmarks, leaderboard shifts, and research from the APEX team.