Research areas and selected projects.
TokenWave works on benchmarks, synthetic data, agents, reinforcement learning, foundation models, and pretraining compute allocation with the M-A-P community.
SuperGPQA
Graduate multiple-choice evaluation across 285 subfields, with sample-weighted and taxonomy-balanced views.
NL2Repo-Bench
Complete Python libraries generated from one requirements document, with correctness decided by the original hidden test suite.
Workflow-GYM
Specialized desktop work scored at the final artifact or GUI state after 30 to 110 dependent actions.
OpenCoder
OpenCoder is a code-language-model and reproducibility paper, not a benchmark.
OProver
OProver is an agentic Lean 4 theorem-proving system whose retrieval, compiler feedback, and proof repair are represented during training as well as inference; it is not a new benchmark.
Ouro
Ouro is a family of 1.4B- and 2.6B-parameter foundation models built with a looped decoder architecture, not an evaluation benchmark.
YuE
YuE is an open foundation-model family for long-form lyrics-to-song generation, not a benchmark.
MAmmoTH2
MAmmoTH2 is primarily a data and instruction-tuning study built around WEBINSTRUCT, a corpus of 10 million instruction-response pairs (about 5 billion tokens) mined from web pretraining data.
Benchmark
Benchmarks for knowledge, coding, agent, and multimodal tasks, with occupation-based task design and automated verification.
Synthetic Data
Methods for mining, generating, and reconstructing instruction data at web scale.
Agents
Systems that plan, execute, and verify multi-step tasks in data science, media production, and other workflows.
Foundation Models
Open text, code, math, and music model families released with data pipelines, training recipes, and checkpoints.
Pretraining Compute Allocation
Looped architectures, dynamic depth, and routing methods that vary compute by token.