Tether's advanced research initiative released QVAC Genesis III on Wednesday, a 191.43-billion-token synthetic dataset designed to train smaller AI models in STEM reasoning and improve their ability to run on laptops, phones and local servers.

According to a statement, the dataset contains 159.6 million documents spanning 19 disciplines across high school, college and professional levels. These include mathematics, physics, chemistry, computer science and medicine.

Tether said the dataset relies on two training methods. The first, Failure Analysis, feeds a smaller model's mistakes back as material, with a larger "teacher" model tracing where the reasoning failed and working through the correct solution.

The other method is Option-Level Reasoning, which explains why a right answer is right and why each alternative is wrong.

Per the statement, Genesis III researchers found that models trained on the dataset outperformed equivalent Cosmopedia-v2 training across several architectures, including Qwen, Llama, SmolLM and Gemma.

Compared with a token-matched Cosmopedia-v2 model, Genesis III improved scores by 28.57 percentage points on ARC-Easy, 21.35 points on ARC-Challenge and 15.03 points on the MMLU STEM benchmark, Tether said.

Genesis III follows Genesis I on Oct. 24, 2025, and Genesis II on Dec. 22, 2025. Its paper has been accepted for presentation at the Conference on Language Modeling 2026.