# Genesis 1B Run 2 Complete - 80,000 Steps - Kroonen AI

> Genesis 1B Run 2 complete: 80,000 steps (~42B tokens) on 2× RTX 4090, final training loss 1.87.

Canonical: https://www.kroonen.ai/blog/genesis-training-progress/

[Lab notebook](https://www.kroonen.ai/blog/) / [Genesis 1B](https://www.kroonen.ai/blog/#genesis) / Entry 006

Run log Run 2 Entry 006

# Genesis 1B: Run 2 Complete at 80,000 Steps

Run 2 complete: 80,000 steps and roughly 42B tokens on two RTX 4090s, with a final training loss of 1.87.

Author

Robin Kroonen

Published

Mar 21, 2026

Updated

Apr 24, 2026

Reading time

8 min read

Steps

80,000

Tokens

~42B

Hardware

2× RTX 4090

Final loss

1.87

Entry 006

Program

Genesis 1B

Format

Run log

Phase

Run 2

Subjects

Genesis / Run 2 / pretraining

[Program index →](https://www.kroonen.ai/blog/#genesis)

Result / Run 2 complete: step 80,004 / 80,000, final loss 1.873

Run 2 reached its 80,000-step target (~42B tokens) on 2× RTX 4090. Explore the checkpoints in the live playground below.

Table of Contents

1.  [Model: Genesis 1B](#model-specs)
2.  [Training Configuration](#training-config)
3.  [Run 2: Progress (40k to 80k) Done](#training-progress)
4.  [The Dataset](#the-dataset)
5.  [The Road to Genesis 1B v0.1](#road-to-v01)
6.  [Run 1: What Happened](#run1-history)
7.  [Try It Yourself](#try-it-yourself)

## Model: Genesis 1B

<table><tbody><tr><td>Parameters</td><td>1,000M (1.0B)</td></tr><tr><td>Architecture</td><td>Llama-style decoder-only transformer</td></tr><tr><td>Hidden dim</td><td>1536</td></tr><tr><td>Layers</td><td>32</td></tr><tr><td>Attention heads</td><td>12 (6 KV heads, GQA)</td></tr><tr><td>FFN dim</td><td>4736 (SwiGLU)</td></tr><tr><td>Context length</td><td>2048</td></tr><tr><td>Vocab size</td><td>49,152</td></tr><tr><td>Precision</td><td>bfloat16</td></tr><tr><td>Positional encoding</td><td>RoPE (θ=500,000)</td></tr></tbody></table>

## Training Configuration

<table><tbody><tr><td>GPUs</td><td>2× RTX 4090 (PCIe, no NVLink)</td></tr><tr><td>Batch size</td><td>4 per GPU</td></tr><tr><td>Gradient accumulation</td><td>32 steps</td></tr><tr><td>Effective batch</td><td>524,288 tokens/step</td></tr><tr><td>Learning rate</td><td>1e-4 → 1e-5 (cosine decay)</td></tr><tr><td>Warmup</td><td>1,000 steps</td></tr><tr><td>Optimizer</td><td>AdamW (β1=0.9, β2=0.95, wd=0.1)</td></tr><tr><td>Activation checkpointing</td><td>Enabled (per TransformerBlock)</td></tr><tr><td>DCP resume</td><td>ShardedStateDictConfig(offload_to_cpu=True)</td></tr><tr><td>CUDA allocator</td><td>expandable_segments:True</td></tr><tr><td>VRAM per GPU</td><td>~20 GB with activation checkpointing</td></tr><tr><td>Throughput</td><td>~19,000 tok/s</td></tr><tr><td>Target</td><td>~42B tokens (80,000 steps, extended from 40,000)</td></tr><tr><td>Script</td><td><code>pretrainv3.py</code></td></tr><tr><td>NCCL</td><td>NCCL_P2P_DISABLE=1</td></tr></tbody></table>

## Run 2: Training Progress (20k → 80k Extension)

Run 2 launched March 24, 2026 with a redesigned 32-layer architecture and reached **20,000 steps** on March 31, 2026. The run was extended first to **40,000 steps** (~21B tokens), completing April 7, 2026 with loss ~1.93, then to **60,000 steps** (~31.5B tokens), and finally to its **80,000-step** target (~42B tokens). Run 2 finished at step **80,004** with a final training loss of **1.87**, throughput holding steady at ~19,000 tok/s throughout.

| Step | Loss | Grad Norm | tok/s |
| --- | --- | --- | --- |
| 0 | 11.1377 | 20.00 | 17,425 |
| 1,000 | 3.4161 | 0.74 | 18,936 |
| 2,000 | 3.0866 | 0.30 | 18,954 |
| 3,000 | 2.5517 | 0.22 | 18,948 |
| 4,000 | 2.6568 | 0.22 | 18,958 |
| 5,000 | 2.2971 | 0.17 | 18,946 |
| 6,000 | 2.2877 | 0.18 | 18,935 |
| 7,000 | 2.2235 | 0.17 | 18,936 |
| 8,000 | 2.1325 | 0.16 | 18,947 |
| 9,000 | 2.2878 | 0.16 | 18,830 |
| 10,000 | 2.1776 | 0.16 | 18,955 |
| 11,000 | 2.1164 | 0.16 | 18,960 |
| 12,000 | 2.2426 | 0.16 | 18,967 |
| 13,000 | 2.1838 | 0.16 | 18,971 |
| 14,000 | 2.0864 | 0.17 | 18,978 |
| 15,000 | 1.9520 | 0.17 | 18,975 |
| 16,000 | 1.8105 | 0.15 | 18,965 |
| 17,000 | 2.1301 | 0.16 | 18,956 |
| 18,000 | 2.1521 | 0.18 | 18,869 |
| 19,000 | 1.8729 | 0.16 | 18,973 |
| 20,000 | 2.2103 | 0.17 | 17,228 |
| 25,000 | 1.9375 | 0.17 | 18,910 |
| 30,000 | 1.9676 | 0.18 | 18,910 |
| 35,000 | 1.9055 | 0.18 | 18,914 |
| **40,000** | **~1.93** | 0.19 | 18,925 |
| 45,000 | 1.8044 | 0.20 | 19,064 |
| 50,000 | 1.8830 | 0.20 | 19,053 |
| 55,000 | 2.0567 | 0.21 | 19,059 |
| **60,000** | **1.9193** | 0.21 | 19,043 |
| **80,004 (final)** | **1.873** | \- | ~19,000 |

Training loss curve

Training loss across Run 2, smoothed; the faint band behind the line is the per-step spread. Final checkpoint: step 80,004, loss 1.87 (~42B tokens on 2× RTX 4090).

At 20k steps, the log shows loss 2.2103 (noisy single-step value). The run continued cosine decay toward 1e-5 through 40k steps. Average loss over steps 38k–40k is **~1.93**. The run was then extended to 60,000 steps across several resumed segments (April 9–17, 2026). Loss at step 60,000 was **1.9193**, throughput held at ~19,050 tok/s throughout. From there the run continued to its 80,000-step target, finishing at step **80,004** with a final training loss of **1.87**.

Checkpoints are backed up locally every 10 minutes. Selected checkpoints can be explored in the [live playground](https://huggingface.co/spaces/rob-x-ai/genesis-1b-run2-playground).

## The Dataset

~60B tokens, curated from public sources:

-   FineWeb-Edu (English web, educational filter)
-   DCLM baseline + extra slices
-   StarCoderData (code)
-   FineMath (mathematics)
-   Wikipedia (multilingual)
-   CulturaX (Arabic, German, Spanish, French, Japanese, Korean, Portuguese, Chinese)
-   OpenHermes, Orca AgentInstruct (instruction data)
-   Function calling datasets (Glaive, Gorilla, Hermes, xLAM)
-   Cosmopedia (synthetic textbooks)

All tokenized with a custom SentencePiece BPE tokenizer trained on the corpus itself.

## The Road to Genesis 1B v0.1

Pre-training is only the first phase. The full pipeline has four stages:

### Phase 1: Pre-training complete (80,000 steps)

Completed at step 80,004 (~42B tokens). Final training loss 1.87.

### Phase 2: SFT (Supervised Fine-Tuning)

SFT runs on top of the pre-trained base. The dataset is 510,577 examples across constitutional data (generated with Claude Haiku), SmolTalk, OpenHermes 2.5, Tulu 3, and MetaMathQA. Training runs for 15,955 steps (1 epoch) at 2e-5 peak learning rate. The approach is inspired by Constitutional AI: define a set of principles and train the model to follow them. The goal is a model with genuine personality, not a model optimized for refusal rates.

### Phase 3: DPO (Direct Preference Optimization)

Refine taste and style. Train the model to prefer interesting, thoughtful responses over generic safe ones. Preference pairs are constructed to reward curiosity and penalize hedging.

### Phase 4: Continued pre-training cycles

Run SFT and DPO on the 40k base, then continue pre-training to 80,000 and beyond. Each cycle produces a better pre-trained foundation, which produces a better aligned model.

The 60B token corpus means zero data repetition even at extended step counts. Every token the model sees is genuinely new data.

## Run 1: What Happened (Historical)

📜 Run 1 History - Click to expand (steps 0-8,500, March 17-24)

Run 1 used a different architecture: 20 layers, dim 2048, 16 heads, batch size 1. It achieved 6,500 tok/s and was on track for ~13 days to 20k steps. Two critical failures occurred:

#### 1\. FSDP Checkpoint Deadlock

Checkpoint saves hung indefinitely due to NCCL ALLGATHER over PCIe without NVLink. Fixed by switching to DCP sharded checkpoints.

#### 2\. Optimizer State Bug (Silent)

The DCP resume path only loaded model weights, not AdamW optimizer state. This produced a false recovery - loss looked healthy for ~50 steps, then diverged. The fix: load optimizer state alongside model weights with try/except fallback.

These failures led to the Run 2 redesign. See the full postmortems: [FSDP Deadlock](https://www.kroonen.ai/blog/genesis-checkpoint-failures/) · [Optimizer State Bug](https://www.kroonen.ai/blog/genesis-optimizer-state-bug/)

#### Run 1 Loss Data

| Step | Loss | Step | Loss |
| --- | --- | --- | --- |
| 0 | 11.17 | 3,400 | 2.73 |
| 200 | 4.87 | 3,600 | 2.42 |
| 400 | 4.34 | 3,800 | 2.45 |
| 600 | 3.55 | 4,000 | 2.25 |
| 800 | 3.03 | 4,200 | 2.35 |
| 1,000 | 3.27 | 4,400 | 2.19 |
| 1,200 | 3.02 | 4,600 | 2.46 |
| 1,400 | 3.02 | 4,800 | 2.10 |
| 1,600 | 2.94 | 5,000 | 2.39 |
| 1,800 | 2.74 | 5,500 | 2.26 |
| 2,000 | 2.54 | 6,000 | 2.20 |
| 2,200 | 2.36 | 6,500 | 2.15 |
| 2,400 | 2.44 | 7,000 | 1.90 |
| 2,600 | 2.54 | 7,500 | 1.69 |
| 2,800 | 2.62 | 8,000 | 1.53 |
| 3,000 | 2.68 | 8,500 | 1.42 |

## Try It Yourself

The model is ready to inspect. Select a checkpoint and generate text to see how it evolved across the run:

Powered by [HuggingFace ZeroGPU](https://huggingface.co/spaces/rob-x-ai/genesis-1b-run2-playground), free inference on NVIDIA H200

Notebook / Reading path

## Continue in Genesis 1B

[Complete index →](https://www.kroonen.ai/blog/)

[

Previous entry 005 / 5 min read

### Genesis 1B, Run 2: Nearly 3× Throughput, Same Hardware

Read entry →](https://www.kroonen.ai/blog/genesis-architecture-v2/)[

Previous entry 004 / 8 min read

### The Optimizer State Bug: A Silent Failure in DCP Resume

Read entry →](https://www.kroonen.ai/blog/genesis-optimizer-state-bug/)

[← Lab notebook](https://www.kroonen.ai/blog/) [Back to top ↑](#main-content)
