Marin: Week of August 17th summary

Milestone: August milestone: launch 535B-A23B MoE
Contents
  1. Data
  2. Summary
  3. 535B-A23B hero launched on 704 B200s; checkpoint reliability rebuilt under load
  4. Snowball SFT checkpoints go public; E6 RLVR reruns end to end with MuonH
  5. Other Changes
  6. Community Pulse
  7. Agent MoE
  8. Runs
GitHub
153 merged 33 opened 73 issues closed 19 contributors 2 epics 267 comments this week
Compute
GCP TPU 4.31e23 HW FLOPs (1.46e23 reserved) W&B 2.67e24 HW FLOPs (4.89e23 model FLOPs)
Compute calculations should be taken with a large grain of salt.
Infra
Discord
286 messages 58 authors 8 new members 22 channels active 22 threads
Tokens
25.6T tokens 0 25.2% synthetic 293 datasets 🤗 collection
web 14.5T (56.5%) code 5.9T (23.2%) multilingual 4.1T (15.8%) specialized 778.1B (3.0%) math 377.2B (1.5%) documents 1.5B (0.0%)

The 535B-A23B hero run launched Tuesday on 704 B200 GPUs across 11 CoreWeave NVL72 racks — the culmination of months of architecture selection, EP64 (64-way expert parallelism) debugging, kernel fusion, and data-mix preparation. A five-rung scaling ladder predicts approximately 2.04 dropless Paloma macro loss at the hero's 2.70e24 training FLOPs. The first days exposed checkpoint-save OOMs recurring every three hours and a restore stall that blocked 703 GPUs for four and a half hours; both were root-caused and fixed within the week, and replica-aware restore cut S3 checkpoint traffic by 57.7%. Two alternative MoE transports — ragged all-to-all and Mixture-of-Kittens — each reached parity with the production pooled-wave backend in live A/B tests from the hero's step-6000 checkpoint, with substantially lower token drop rates.

Meanwhile, the June 67B-A2B "Grug" run on TPU v4-2048, which was at 89% last week, continued training with a new context-extension branch at 262K tokens from step 156k. The full Snowball SFT checkpoint chain went public on Hugging Face, closing a loop that started with the cold-start pipeline delivered the prior week. On the post-training side, the E6 RLVR (reinforcement learning with verifiable rewards) rerun with MuonH reached an end-to-end completion from the Marin graph — the first successful reinforcement learning run through the full artifact pipeline.

Hero Runs

The milestone’s pretraining hero runs and their intermediate cooldowns — the concrete use of compute.

#6689 535B-A23B hero launched on 704 B200s; checkpoint reliability rebuilt under load

Epic title: [Hero Run] ~120B-A8B XT on B200s


Summary: Prepare the next best model for post-training on the path to our EOY 256–500B-AYB run.

0/3 sub-issues closed

The 535B-A23B EP64 (64-way expert parallelism) MoE hero run launched on August 19 across 704 B200 GPUs (11 NVL72 racks on CoreWeave), training on 18 trillion tokens of the Harrier data mix. Larry Dial posted the full model specification—d6144, 384 experts top-8, pooled-wave all-to-all transport—and a five-rung scaling-ladder analysis in #8435: four clean rungs and a d2048 rung that reached 81% yield a power-law prediction of approximately 2.04 dropless Paloma macro-loss at the hero’s 2.70e24 training FLOPs. Will Held finalized the Harrier mixture on a fuzzy-deduplicated datakit store #8427 with monotonic PI (proportional importance) weights #8452, and added uncheatable evals to every ladder rung #8425. Percy Liang announced the sprint in Discord; a dedicated #hero-run-2026 channel opened for public tracking.

The first days hammered the checkpoint pipeline. Host OOM kills recurred roughly every three hours during saves #8506. Russell Power walked through malloc_trim #8540, jemalloc #8550, staging-budget raises from 4 to 16 GiB per process #8514, and write-replica fanout from 64 to 1024 #8486, while Mark Muchane routed checkpoints to cluster-local storage #8559 and hero paths were locked against deletion #8504. A separate restore stall blocked 703 GPUs for four and a half hours when one rank’s TensorStore read hung #8534. Matt Wittmann root-caused it to an unbounded post-restore barrier and landed a timed replacement #8538; Russell Power added a TensorStore low-speed abort #8584. Replica-aware restore now reads each unique checkpoint shard once and distributes via JAX reduction, cutting benchmark S3 traffic by 57.7% #8589. Rafal Wojdyla built a training-run status dashboard #8479, wired three hero-health alert rules through Slack #8568, and pulled the full loss curve from W&B after Finelog evicted early steps #8566.

Two MoE transport alternatives reached parity with the production pooled-wave backend. Matt Wittmann’s ragged EP64 rewrite, tested in a checkpoint-restored A/B from the live hero’s step-6000 state, delivered 22.46 vs 22.52% MFU with 0.015% drops against 2.67% and a 10 GiB lower runtime device peak #8549 #8317. Mark Muchane’s Mixture-of-Kittens implementation matched pooled-wave at EP64 hero shape—237,613 vs 237,310 tokens/s—with zero drops #8108. On the kernel side, David Hall fixed silent FP8 quality loss from gradient accumulation dropping the amax history to the last microbatch #8360 #8365. Larry Dial merged the combined hero-kernel stack: bf16 backward GEMMs in fused cross-entropy (a 2.33x speedup), fused ShortConv, and hoisted router all-reduces #8385. Will Held’s Pallas short-convolution experiments passed Gate 1 at both d512 (1.242x combined speedup) and d768 (1.131x) #8377.

0 PRs this week, and 0 new issues (3 total)
Sort:
117 autocategorized
53 potentially related in Other Changes

#6705 Snowball SFT checkpoints go public; E6 RLVR reruns end to end with MuonH

Epic title: [Hero run] Post training on 67B-A2B 10T


Summary: Listed here for discussion for July planning, realistically the final hero run training on the 67B-A2B on full 10T tokens won't begin until early (or mid?) August: we should have a different issue for the training and debugging on the intermediate smaller-token-count cuts

The full Snowball 67B-A2B cold-start SFT checkpoint chain is now public on Hugging Face. Benjamin Feuer published the base cooldown at step 105149, the Chat stage at step 257, Thinking at step 630, and both Agentic branches (OpenCode and Nemotron Terminal) at steps 1000 and 1888, covering every durable checkpoint the training runs retained. A community request for intermediate post-training checkpoints #8341 was resolved by the release. On the source-run side, Larry Dial reported that the 67B-A2B at 10T tokens has reached Marin’s lowest Paloma macro loss ever, noting the comparison is not strict since this phase runs at 65k context length. The loss slope accelerated after entering mix2 and again on longer context, with 2T tokens now on mix2 and 1T of that on longer context. A W&B report tracks the full run.

Ahmad Qamar brought the Snowball E6 RLVR (reinforcement learning with verifiable rewards) rerun with MuonH to an end-to-end completion from the Marin graph. MuonH reached a learning signal where the earlier arm E10-B had been blocked on OOM around step 5. The run used a reduced geometry (5 nodes, batch size 32, 6 steps on CoreWeave) due to capacity constraints, so reward and throughput are not directly comparable to the original #7786. Four fixes were required: #8509 quoted Hydra retention overrides so the RL stage no longer died at launch, #8510 surfaced launcher stderr in error messages, #8565 added ExternalModel and ExternalDataSource so SkyRLSpec can reference published checkpoints and datasets, and MarinSkyRL#423 fixed job-status recording. Benjamin Feuer flagged a learning-rate footgun in prior attempts: Muon-H matrices and the language-model head used the configured 1e-5 rate while embeddings, router weights, and norms fell through to Adam at an implicit 6e-4, which MarinSkyRL PR #405 corrected.

Separately, Qamar opened a coordinated PR set for RL observability and identity: #8562 and #8551 add a dedicated Grafana dashboard for RL runs (the existing one filters only Levanter pretraining), and #8561 and #8555 derive run IDs from the artifact’s storage address instead of a version basename, so that an SFT and a pretraining checkpoint at the same version no longer collide. On the data-selection front, Feuer published three experiment results informing what to train on: a 48-cell harness/context/summarization ablation finding that harness preference is model- and dataset-dependent, a cross-cluster HPO (hyperparameter optimization) study showing a wide tolerance band for RL hyperparameters but that HPO matters for long-run stability, and a five-model reward-delta analysis identifying 24+ TaskTrove sources with monotonic, substantive gaps across Qwen3 Coder, Qwen3.5 122B, and GLM 5.2. A culminating five-model data quality panel on TaskTrove v3.42 also found that model capacity is not predictive of agentic benchmark performance.

22 autocategorized
10 potentially related in Other Changes

Other Changes


Iris received a wave of reliability and operability work. Russell Power and agents landed Kubernetes disruption evidence retention (#8601, #8604), a fix for kubelet resource metric decoding after the Kubernetes client v36 upgrade broke Prometheus parsing (#8582, regression from #8381), and grace-period handling for transient unknown pod phases so the hero run stops burning failure quota on ephemeral Kubernetes hiccups #8587. #8463 fixed a bug where Iris’s callable runner swallowed fatal exceptions, letting the clean-exit hook mark failed jobs as succeeded. The workload client and CLI were normalized: #8364 and #8399 expose immutable Job/Task/Attempt handles with a unified verb vocabulary (job cancel, job complete, task preempt), replacing ad hoc commands. #8424 preserves terminal causes like OOMKilled through job summaries. Ryan Williams re-landed the client-freshness gate (#8522, #8523), anchoring the version floor to the controller’s own build so quiet weeks no longer reject clients running identical code. Romain Yon upgraded the Kubernetes client to v36.0.3 #8381 and excluded GitHub Actions credentials from workspace bundles #8518.

Observability improved across Grafana, Finelog, and fleet telemetry. Russell Power cut Finelog query latency by orders of magnitude: structured telemetry queries dropped from 8.5 seconds to 200 milliseconds through index bypass and timestamp-bound planning #8379, shared-prefix IN queries now prune with a half-open range #8516, and Cluster Capacity dashboard medians fell from 838 to 244 milliseconds #8526. A new per-GPU SM (streaming multiprocessor) utilization raster replaced the fleet histogram #8513, and a cluster capacity packing view now shows live pod placement against node capacity #8423. Rafal Wojdyla exposed Levanter training status through Iris #8512 with ad hoc checkpoint support (#8544 in the hero epic), and the log viewer gained paging, time-bound controls, and bulk context expansion #8493. Finelog Kubernetes deploys moved to Pulumi #7690, contributed by Will Moss. The Evalchemy and Harbor nightly dashboards stopped flagging fast successful runs as unhealthy (#8521, #8532).

Infrastructure-as-code and storage management saw substantial changes. Russell Power moved shared GCS, CoreWeave, and R2 bucket management into a dedicated Pulumi stack #8343, replacing the old configure_buckets.py script. CoreWeave bucket policies were scoped to the Open Athena organization while denying object deletion under hero checkpoint prefixes #8588. The distributed storage scanner was generalized to cover GCS, CoreWeave, and R2 backends #8382, and fsutil rm now streams deletes while listing rather than materializing the full object set first #8355. The synthetic infrastructure probe was retired after all 38 runs since July were cancelled at timeout #8473, and the flaky GCP pull-request smoke test was removed after a 33% failure rate from TPU capacity issues #8408. Canary runs switched to a fixed 10M-parameter model with one-day outputs (#8600, #8545) and region-local placement #8605.

The data pipeline gained several fixes. Rafal Wojdyla fixed Zephyr’s object-store client leak: each shard had been creating a fresh S3 client, parking connections in TIME_WAIT until port exhaustion killed the process #8406. Zephyr per-stage counter totals were fixed #8353, CoreWeave spill reads were corrected for LOTA path-style rewriting #8482, and Zephyr’s InputFileSpec.format field, previously write-only, now actually controls the file reader #8530. FineStore gained bounded-part storage for large objects #8466. Zephyr A/B benchmarks defaulted to GCP (#8342, #8451).

Developer tooling improved across the board. Romain Yon indexed six community repositories in Echo #8448, with per-repository scoping so cross-repo search does not dilute results. EvalDash now serves from a PostgreSQL catalog with a 3.5-second startup instead of 3.5 minutes #8346. Loom sessions now submit Iris jobs as user loom instead of app #8488, and Slack sessions default to medium reasoning effort #8433. Agent skill context was trimmed by 15% #8494. Benjamin Feuer fixed Evalchemy to exclude infrastructure-error completions from task scores #8337. David Hall published the Agent MoE experiment digest covering 80 tracker experiments #7623 and fixed FP8 state accumulation across microbatches (#8360, #8365). The Whisper model’s cross-attention initialization bug, which gave cross-attention projections bit-identical weights to self-attention, was also fixed #8446.

120 PRs this week, 93 new comments, and 73 issues closed (73 total)
Sort:

Community Pulse


Will Moss (Airbnb) opened #8497 adding A/B test benchmark sizing guidance for the Datakit pipeline, and a new contributor Ian Morgan submitted #8591 bounding CI system-dependency setup so slow mirrors cannot consume test budgets.

In #news, Colin Raffel, Will Held, and Jenia Jitsev debated a paper on small-scale experiment reliability for scaling laws — Will Held contrasted it with Marin’s Tensor Programs-style heuristics, while Colin Raffel noted the paper’s finding that hyperparameter sensitivity decreases at scale. In #questions, Jenia Jitsev raised the question of Marin model adoption in vLLM and HuggingFace; Percy Liang outlined a phased plan, and Romain Yon explained why Marin currently uses its own Levanter architecture for experimental freedom. In #moe, catto proposed an initialization trick to guarantee perfect load balancing on the first batch without the QB solve; Larry Dial said they would test it for future runs.

Eight people introduced themselves. Nobin Sarwar (PhD, UMBC; scientific reasoning and LLM safety), Carlos Aspillaga (CENIA, Chile; low-resource languages), Jay Zhou (PhD, USC), Abel (CS graduate, Universitas Indonesia), Evan Quintana (research intern, Sandia National Labs; AI security), and a research engineer named Alex focused on kernels. Leonard, co-founder and CTO of Living Models, joined alongside Jean du Terrail to build foundation models for living systems — a direct connection to the MarinFold protein-structure work, where Gonzalo Benegas welcomed the overlap. Lena Lincke, starting a PhD at TU Munich with a classical RL background, is focused on LLM post-training — the area where the team’s next hero run lands after pretraining.

Shared research centered on MoE scaling: a hyperparameter transfer study for a 155B-A17B MoE at 10T tokens, and a PrimeIntellect thread on the current limits of agentic research.

News & research shared

Active collaborators this week

Stanford · CRFM 5 people · 26 Discord msgs

Collaborator activity this week

Lab / Org People PRs Issues filed Comments Discord msgs Total
Stanford · CRFM 5 26 26
CMU · NeuLab
Common Crawl Foundation
Princeton · Dao Lab
GitHub activity from 50 other contributors

Tim O'Donnell · McGill · (other) 0 PRs, 34 Discord msgs

Will Moss · Industry (other) 1 PR, 10 comments

  • #8497 [zephyr][datakit] Provide guidance on A/B test benchmark sizing 💬1 +506 −42
10 comments on 9 threads
  • #7580 [pulumi] Manage storage buckets, lifecycle rules, and bucket-scoped IAM grants in Pulumi ×2
  • #7690 [finelog] Move Kubernetes deploys to Pulumi
  • #8497 [zephyr][datakit] Provide guidance on A/B test benchmark sizing
  • #8458 [iac] Centralize application IAM grants
  • #7906 [pulumi] Detect drift between Pulumi and live GCP infrastructure
  • #8363 [iac] Centralize GCS bucket access
  • #7713 [pulumi] Periodically scan for GCP IAM drift outside Pulumi
  • #8455 [pulumi] Make IAM grants authoritative
  • #8457 [weaverbot] Testing

Mrinal Kumar · Unclassified 0 PRs, 13 Discord msgs

boba shop cashier · Unclassified 0 PRs, 10 Discord msgs

lukedhlee · Unclassified 0 PRs, 10 Discord msgs

Marianna Nezhurina 0 PRs, 7 Discord msgs

Ayush Sunil Munot · Unclassified 0 PRs, 3 comments, 3 Discord msgs

3 comments on 2 threads
  • #8357 evaluating whether a model spends the right amount of thinking ×2
  • #7090 Epic: new evals for Marin — wish-list

Matheart · Unclassified 0 PRs, 6 Discord msgs

Mayank 0 PRs, 5 Discord msgs

Fluf · Unclassified 0 PRs, 5 Discord msgs

catto · Unclassified 0 PRs, 4 Discord msgs

Jenia Jitsev · LAION 0 PRs, 4 Discord msgs

Bilibird · Unclassified 0 PRs, 3 Discord msgs

Chloe Chia · Independent 0 PRs, 3 Discord msgs

Neha Hulkund · MIT · Open-Thoughts Next 0 PRs, 3 Discord msgs

Colin Raffel · UNC · (other) 0 PRs, 3 Discord msgs

Elman Mansimov · Industry (other) 0 PRs, 1 comment, 1 Discord msg

1 comment on 1 thread
  • #8435 [Hero Run] 535B-A23B on 18T tokens

Rohith Kuditipudi · Stanford · (other) 0 PRs, 1 Discord msg

Nobin Sarwar · Unclassified 0 PRs, 2 Discord msgs

Carlos Aspillaga · Unclassified 0 PRs, 2 Discord msgs

Lena Lincke · TU Munich · (other) 0 PRs, 2 Discord msgs

Leonard · Industry (other) 0 PRs, 2 Discord msgs

Abel · Unclassified 0 PRs, 2 Discord msgs

cdd · Unclassified 0 PRs, 2 Discord msgs

Ian Morgan · Unclassified 1 PR

  • #8591 [ci] Bound system dependency setup in special-test jobs +28 −5

NivC · Unclassified 0 PRs, 1 comment

1 comment on 1 thread
  • #8341 Intermediate checkpoints during post-training

Yiyuan Li · UNC · (other) 0 PRs, 1 comment

1 comment on 1 thread
  • #8357 evaluating whether a model spends the right amount of thinking

Windsor Nguyen · Industry (other) 0 PRs, 1 comment

1 comment on 1 thread
  • #8435 [Hero Run] 535B-A23B on 18T tokens

Michael Siu · Unclassified 0 PRs

Cerise · Unclassified 0 PRs

Tinuade Adeleke 0 PRs, 1 comment

1 comment on 1 thread
  • #7617 [data] Fix the focus-crawl PDF input and the output shape

Dmitrii · Unclassified 0 PRs, 1 Discord msg

charlene · Unclassified 0 PRs, 1 Discord msg

Adyasha · Unclassified 0 PRs, 1 Discord msg

boom · Unclassified 0 PRs, 1 Discord msg

howard · Unclassified 0 PRs, 1 Discord msg

jan_s · Unclassified 0 PRs, 1 Discord msg

Joy Jing · Unclassified 0 PRs, 1 Discord msg

PUchoose · Unclassified 0 PRs, 1 Discord msg

alex · Unclassified 0 PRs, 1 Discord msg

Jay Zhou · Unclassified 0 PRs, 1 Discord msg

Evan Quintana · Unclassified 0 PRs, 1 Discord msg

Luca · Unclassified 0 PRs, 1 Discord msg

omar_midjourney · Unclassified 0 PRs, 1 Discord msg

nato · Unclassified 0 PRs, 1 Discord msg

Charlie R · Unclassified 0 PRs, 1 Discord msg

Jay Ning · Unclassified 0 PRs, 1 Discord msg

jku100 · Unclassified 0 PRs, 1 Discord msg

elie · Unclassified 0 PRs, 1 Discord msg

samsja · Unclassified 0 PRs, 1 Discord msg

Agent MoE speedup


Completed marin-community/marin_moe runs, grouped by Agent MoE budget. Speedup is relative to the original baseline run for each budget and charges each variant by its actual reported FLOPs. Best observed point is 17.40× from mhep-ladder-hist-20260808c-fsdp-chunk1-d1024.

baseline (1×) this week's runs older runs running best higher is better
d512 / 2.19e17 FLOPs
100 completed runs; 2 this period
baseline loss 3.8104
1x moe_may_compute_opt_d512_ep1_baseline: 0.79x, loss 3.5448, Jun 20 moe_may_compute_opt_d512_ep1_normswish: 0.79x, loss 3.5448, Jun 20 moe_may_compute_opt_mla_d512: 0.50x, loss 3.6619, Jun 20 moe_may_compute_opt_gqa_d512: 0.55x, loss 3.6555, Jun 20 moe_may_compute_opt_mla_norm_compressed_d512: 0.46x, loss 3.6811, Jun 20 moe_may_compute_opt_d512_ep1_normswish_vector: 0.80x, loss 3.5428, Jun 20 moe_may_compute_opt_d512_ep1_normswish_scalar: 0.79x, loss 3.5437, Jun 20 june_prep_moe_may_d512_ep2_no_long_rope_64kctx_mscale1p1_from71808: 0.05x, loss 3.1350, Jun 21 june_prep_moe_may_d512_ep2_no_long_rope_64kctx_mscale1p3_from71808: 0.05x, loss 3.1336, Jun 21 june_prep_moe_may_d512_ep2_no_long_rope_64kctx_mscale1p0_from71808: 0.05x, loss 3.1362, Jun 21 moe_may_compute_opt_d512_validate_seq4k_v5p8_fp32ns: 0.78x, loss 3.5472, Jun 25 moe_may_compute_opt_d512_validate_seq4k_v5p8: 0.79x, loss 3.5487, Jun 25 moe_may_5000tn_4x_d512_ep2_v1_adamh_warmup1pct_e256: 0.23x, loss 3.2103, Jun 25 moe_may_5000tn_4x_d512_ep2_v1_adamh_warmup1pct_e256_v2: 0.24x, loss 3.2042, Jun 25 moe_compute_opt_d512_stacked_rmsadam_v5p_8: 0.71x, loss 3.5711, Jun 27 moe_compute_opt_d512_stacked_baseline_v5p_8: 0.72x, loss 3.5676, Jun 27 swarm_fisher_dsp_d512_000850: 0.06x, loss 3.3087, Jul 8 grug-copt-d512-evalfix-20260709-015252: 1.36x, loss 3.7028, Jul 9 grug-copt-d512-e256-evalfix-20260709-024801: 1.45x, loss 3.6494, Jul 9 grug-copt-d512-e256-nosim-sharedH-20260709-035332: 1.29x, loss 3.6081, Jul 9 grug-copt-d512-e256-pko-longrope-20260709-044728: 1.43x, loss 3.6294, Jul 9 grug-copt-d512-e256-pko-vmap3d-20260709-060901: 1.24x, loss 3.6150, Jul 9 swarm_fisher_dsp_d512_000851: 0.06x, loss 3.3118, Jul 9 swarm_fisher_dsp_d512_000853: 0.06x, loss 3.3119, Jul 9 swarm_fisher_dsp_d512_000857: 0.06x, loss 3.3091, Jul 9 grug-mainstack-d512-copt-20260709-144031: 1.03x, loss 3.7103, Jul 9 grug-mainstack-vmap-d512-copt-20260709-144127: 1.16x, loss 3.6954, Jul 9 swarm_fisher_dsp_d512_000848: 0.06x, loss 3.3082, Jul 9 grug-mainstack-vmap-d512-e256-copt-20260709-152624: 1.02x, loss 3.6359, Jul 9 swarm_fisher_dsp_d512_000847: 0.06x, loss 3.3085, Jul 9 swarm_fisher_dsp_d512_000846: 0.06x, loss 3.3138, Jul 9 swarm_fisher_dsp_d512_000858: 0.06x, loss 3.3118, Jul 9 swarm_fisher_dsp_d512_000856: 0.06x, loss 3.3108, Jul 9 swarm_fisher_dsp_d512_000849: 0.06x, loss 3.3090, Jul 9 swarm_fisher_dsp_d512_000862: 0.07x, loss 3.2998, Jul 9 grug-tpu-v5p8-d512-e256-sw2048-nemotron-pko-longrope-copt-20260709-163035: 0.76x, loss 3.5495, Jul 10 grug-tpu-v5p8-d512-e256-sw2048-nemotron-copt-20260709-163106: 0.66x, loss 3.5765, Jul 10 swarm_fisher_dsp_d512_000861: 0.07x, loss 3.2996, Jul 10 swarm_fisher_dsp_d512_000863: 0.06x, loss 3.3062, Jul 10 grug-tpu-v5p8-d512-e256-copt-20260709-151454: 0.45x, loss 3.6469, Jul 10 grug-tpu-v5p8-d512-e256-sw2048-nemotron-pko-longrope-minlr0-copt-20260709-215032: 0.79x, loss 3.5421, Jul 10 grug-tpu-v5p8-d512-e256-sw2048-nemotron-pko-longrope-minlr0-evalf32-copt-20260709-223745: 0.79x, loss 3.5420, Jul 10 grug-tpu-v5p8-d512-e256-sw2048-copt-20260709-162033: 0.47x, loss 3.6403, Jul 10 swarm_fisher_dsp_d512_000872: 0.06x, loss 3.3119, Jul 10 swarm_fisher_dsp_d512_000877: 0.06x, loss 3.3144, Jul 10 swarm_fisher_dsp_d512_000883: 0.06x, loss 3.3104, Jul 10 swarm_fisher_dsp_d512_000865: 0.06x, loss 3.3102, Jul 10 swarm_fisher_dsp_d512_000893: 0.06x, loss 3.3122, Jul 10 swarm_fisher_dsp_d512_000895: 0.06x, loss 3.3121, Jul 10 swarm_fisher_dsp_d512_000876: 0.06x, loss 3.3114, Jul 10 swarm_fisher_dsp_d512_000894: 0.06x, loss 3.3111, Jul 10 swarm_fisher_dsp_d512_000898: 0.06x, loss 3.3119, Jul 10 swarm_fisher_dsp_d512_000899: 0.06x, loss 3.3108, Jul 10 swarm_fisher_dsp_d512_000896: 0.06x, loss 3.3106, Jul 10 swarm_fisher_dsp_d512_000892: 0.06x, loss 3.3106, Jul 10 swarm_fisher_dsp_d512_000879: 0.06x, loss 3.3085, Jul 10 swarm_fisher_dsp_d512_000891: 0.06x, loss 3.3107, Jul 10 swarm_fisher_dsp_d512_000878: 0.06x, loss 3.3075, Jul 10 swarm_fisher_dsp_d512_000873: 0.06x, loss 3.3099, Jul 10 swarm_fisher_dsp_d512_000868: 0.06x, loss 3.3139, Jul 10 swarm_fisher_dsp_d512_000880: 0.06x, loss 3.3095, Jul 10 swarm_fisher_dsp_d512_000885: 0.06x, loss 3.3107, Jul 10 swarm_fisher_dsp_d512_000887: 0.06x, loss 3.3105, Jul 10 swarm_fisher_dsp_d512_000871: 0.07x, loss 3.3045, Jul 10 swarm_fisher_dsp_d512_000900: 0.06x, loss 3.3096, Jul 12 MOE-MRCR-001-d512-r6: 0.57x, loss 3.6643, Jul 15 aug-hero-d512-60x-lr1-v6: 1.92x, loss 3.6653, Aug 1 aug-hero-d512-30x-lr0.7: 1.11x, loss 3.9087, Aug 1 aug-hero-d512-30x-lr0.85: 1.44x, loss 3.8673, Aug 1 aug-hero-d512-30x-lr1.2: 1.54x, loss 3.8420, Aug 1 aug-hero-d512-30x-lr1: 1.48x, loss 3.8518, Aug 1 aug-hero-d512-60x-lr1.4: 1.46x, loss 3.7107, Aug 1 aug-hero-d512-30x-lr0.7-v2: 1.12x, loss 3.9083, Aug 1 aug-hero-d512-30x-lr1.2-v2: 1.52x, loss 3.8444, Aug 1 aug-hero-d512-30x-lr1.4-v2: 1.52x, loss 3.8469, Aug 1 aug-hero-d512-30x-lr1-v2: 1.50x, loss 3.8478, Aug 1 aug-hero-d512-60x-lr0.7-v2: 1.63x, loss 3.6919, Aug 1 aug-hero-d512-60x-lr1-v2: 1.88x, loss 3.6662, Aug 1 aug-hero-d512-60x-lr0.85-v2: 1.82x, loss 3.6714, Aug 1 aug-hero-d512-60x-lr1.4-v2: 1.78x, loss 3.6786, Aug 1 aug-hero-d512-300x-lr0.7-v2: 1.45x, loss 3.4215, Aug 1 aug-hero-d512-300x-lr1.4-v2: 1.46x, loss 3.4157, Aug 1 aug-hero-d512-300x-lr0.85-v2: 1.52x, loss 3.4138, Aug 1 aug-hero-d512-300x-lr1.2-v2: 1.56x, loss 3.4118, Aug 1 aug-hero-d512-300x-lr1-v2: 1.53x, loss 3.4093, Aug 1 iso-1e18-d512: 1.55x, loss 3.3990, Aug 6 grug_xem_d512_smoke_pairwise_sqrt: 0.02x, loss 5.2199, Aug 6 grug_xem_d512_smoke_pairwise_linear: 0.02x, loss 5.2409, Aug 6 grug_xem_d512_smoke_pairwise_unscaled: 0.03x, loss 5.1927, Aug 6 grug_xem_d512_smoke_baseline: 0.03x, loss 5.1615, Aug 6 grug_xem_d512_smoke_middle4_sqrt: 0.02x, loss 5.2454, Aug 6 grug_xem_d512_smoke_middle4_linear: 0.02x, loss 5.2641, Aug 6 grug_xem_d512_smoke_middle4_unscaled: 0.03x, loss 5.2025, Aug 6 grug_xem_d512_full_pairwise_unscaled: 0.59x, loss 3.6067, Aug 6 grug_xem_d512_full_pairwise_sqrt: 0.58x, loss 3.6083, Aug 6 grug_xem_d512_full_baseline: 0.64x, loss 3.5862, Aug 6 grug_xem_d512_full_middle4_sqrt: 0.52x, loss 3.6316, Aug 6 grug_xem_d512_full_middle4_unscaled: 0.54x, loss 3.6244, Aug 6 MOE-PSC-101-d512-v5p8-gate1: 0.65x, loss 3.6227, Aug 18 MOE-PSC-CTRL-101-d512-v5p8: 0.52x, loss 3.6760, Aug 19 Jun 20 Aug 19
Best
1.92× aug-hero-d512-60x-lr1-v6 loss 3.6653
This week
0.65× MOE-PSC-101-d512-v5p8-gate1 loss 3.6227
Baseline
moe-v16-compute-opt-d512-2.19e+17
d768 / 1.70e18 FLOPs
100 completed runs; 10 this period
baseline loss 3.4339
1x mhep-abl-d768-ep-20260810f: 4.00x, loss 3.2400, Aug 11 mhep-abl-d768-baseline-20260812: 3.99x, loss 3.2269, Aug 12 mhep-abl-d768-lnpost-20260812: 3.53x, loss 3.2438, Aug 12 mhep-abl-d768-baseline-bf16grad-20260812: 3.93x, loss 3.2278, Aug 12 mhep-abl-d768-lnpre-matchvar-20260812: 3.98x, loss 3.2274, Aug 12 mhep-abl-d768-lnboth-20260812: 3.78x, loss 3.2367, Aug 12 mhep-abl-d768-shared1-20260812: 3.84x, loss 3.2321, Aug 12 mhep-abl-d768-muon8-polar-20260812: 3.78x, loss 3.2292, Aug 12 mhep-abl-d768-beta1-090-20260812: 3.91x, loss 3.2273, Aug 12 mhep-abl-d768-beta1-091-20260812: 3.92x, loss 3.2268, Aug 12 mhep-abl-d768-e384-t8-half-20260812: 3.70x, loss 3.2219, Aug 12 mhep-abl-d768-e384-t8-mfu-20260812: 0.00x, loss 7.3908, Aug 12 mhep-abl-d768-e384-t8-half-smatch-20260812: 3.57x, loss 3.2256, Aug 12 mhep-abl-d768-e768-t16-quarter-smatch-20260812: 2.83x, loss 3.2253, Aug 12 mhep-abl-d768-baseline-normuon-20260812: 3.81x, loss 3.2325, Aug 13 mhep-abl-d768-pe5-20260812: 3.87x, loss 3.2320, Aug 13 mhep-abl-d768-beta1-092-20260812: 3.93x, loss 3.2276, Aug 13 mhep-abl-d768-beta1-093-20260812: 3.92x, loss 3.2314, Aug 13 mhep-abl-d768-cw4-20260812: 2.52x, loss 3.2280, Aug 13 mhep-abl-d768-cw15-20260812: 2.69x, loss 3.2305, Aug 13 mhep-abl-d768-splitsq-20260812: 2.67x, loss 3.2315, Aug 13 mhep-abl-d768-splitsq-gateup-20260812: 3.79x, loss 3.2306, Aug 13 mhep-abl-d768-splitsq-alllatent-20260812: 4.02x, loss 3.2280, Aug 13 mhep-abl-d768-sconvgate-20260812: 3.86x, loss 3.2303, Aug 13 mhep-abl6711-d768-baseline-20260813: 4.22x, loss 3.2172, Aug 13 mhep-abl6711-d768-no-attn-gate-20260813: 3.92x, loss 3.2303, Aug 13 mhep-abl6711-d768-full-rope-20260813: 3.78x, loss 3.2362, Aug 13 mhep-abl6711-d768-no-qk-norm-20260813: 3.84x, loss 3.2312, Aug 13 mhep-abl6711-d768-no-qk-mult-20260813: 3.99x, loss 3.2256, Aug 13 harrier-proportional-d768-gb200-val-smoke-fix-20260813-1830: 0.00x, loss 11.8170, Aug 14 mhep-d768-baseline-skipc01-marinpath-20260814: 4.53x, loss 3.2008, Aug 14 mhep-abl6711-d768-no-latent-norm-20260813-r5: 3.39x, loss 3.2637, Aug 14 datakit-store-proportional-d768-full-krr-20260813: 4.64x, loss 3.2141, Aug 14 mhep-abl6711-d768-no-xsa-20260814-r6: 4.62x, loss 3.2179, Aug 14 mhep-abl6711-d768-no-gated-norm-20260814-r6: 4.40x, loss 3.2240, Aug 14 mhep-abl6711-d768-no-renorm-20260814-r6: 4.20x, loss 3.2265, Aug 14 mhep-abl6711-d768-adamh-matrices-20260814-r6: 4.28x, loss 3.2226, Aug 14 mhep-abl6711-d768-window-512-20260814-r6: 4.67x, loss 3.2298, Aug 14 mhep-abl6711-d768-no-sconv-20260814-r6: 3.70x, loss 3.2508, Aug 14 mhep-abl6711-d768-mha-20260814-r6: 5.24x, loss 3.1870, Aug 14 harrier-nemotron-proportional-d768-full-krr-20260814: 10.26x, loss 3.1007, Aug 14 mhep-abl6711-d768-sconv-global-20260814-r6: 4.26x, loss 3.2253, Aug 14 mhep-abl6711-d768-qk-gain-20260814-r7: 4.50x, loss 3.2184, Aug 14 harrier-marin-proportional-d768-buffer128-full-krr-20260814: 3.17x, loss 3.2769, Aug 14 harrier-marin-proportional-d768-buffer128-prefetch64-seed1-full-krr-20260814: 2.93x, loss 3.2819, Aug 15 harrier-marin-proportional-d768-buffer256-prefetch128-seed2-full-krr-20260814: 2.94x, loss 3.2842, Aug 15 mhep-d768-offload-bf16-hostmaster-test-20260814: 4.35x, loss 3.2202, Aug 15 mhep-d768-base-fp32-masteroff-test-20260814: 4.16x, loss 3.2219, Aug 15 harrier-marin-proportional-d768-buffer256-initial256-prefetch32-seed3-full-krr-20260814: 2.99x, loss 3.2799, Aug 15 harrier-marin-proportional-d768-buffer256-prefetch256-seed4-full-krr-20260814: 3.03x, loss 3.2835, Aug 15 mhep-ep8062-d768-pool-w3-cf133-20260814: 5.50x, loss 3.1539, Aug 15 harrier-marin-proportional-d768-buffer512-prefetch256-seed5-full-krr-20260814: 3.09x, loss 3.2795, Aug 15 harrier-marin-proportional-d768-buffer512-initial1-prefetch256-seed6-full-krr-20260814: 3.14x, loss 3.2784, Aug 15 harrier-marin-proportional-d768-buffer256-ramp1-32-64-128-256-seed7-full-krr-20260814: 2.74x, loss 3.2933, Aug 15 harrier-marin-proportional-d768-buffer256-ramp8-16-32-64-128-256-seed8-full-krr-20260814: 2.93x, loss 3.2855, Aug 15 harrier-marin-proportional-d768-buffer256-prefetch128-tscache8g-seed9-full-krr-20260814: 3.11x, loss 3.2821, Aug 15 mhep-abl6711-d768-xsa-local-sconv-global-20260814: 4.38x, loss 3.2233, Aug 15 mhep-abl6711-d768-xsa-raw-value-20260814: 4.54x, loss 3.2181, Aug 15 mhep-abl6711-d768-no-sconv-v-20260814: 4.44x, loss 3.2212, Aug 15 mhep-ep8062-d768-pool-w3-cf133-halfexp-20260814: 4.84x, loss 3.1946, Aug 15 mhep-abl6711-d768-xsa-local-vsconv-global-20260814: 4.55x, loss 3.2200, Aug 15 mhep-abl6711-d768-no-xsa-no-sconv-20260814: 3.87x, loss 3.2474, Aug 15 mhep-ep8062-d768-pool-w3-cf115-halfexp-dropsplit-20260814: 3.59x, loss 3.2441, Aug 15 mhep-abl6711-d768-one-shared-20260814-r6: 4.46x, loss 3.2205, Aug 15 mhep-abl6711-d768-sconv-k-attn-20260815-r2: 4.13x, loss 3.2307, Aug 15 mhep-abl6711-d768-sconv-kernel2-nov-20260815: 4.28x, loss 3.2285, Aug 15 mhep-abl6711-d768-baseline-20260815: 4.57x, loss 3.2168, Aug 15 mhep-abl6711-d768-sconv-k-only-20260815-r2: 3.99x, loss 3.2431, Aug 16 mhep-abl6711-d768-sconv-k4-attn2-mlp2-nov-20260815-r2: 4.12x, loss 3.2255, Aug 16 mhep-abl6711-d768-sconv-k4-attn1-mlp1-nov-20260815: 4.04x, loss 3.2374, Aug 16 mhep-abl6711-d768-sconv-mlp-routed-only-20260815-r2: 4.47x, loss 3.2202, Aug 16 mhep-abl6711-d768-no-zloss-20260814-r6: 0.43x, loss 3.6160, Aug 16 mhep-abl6711-d768-sconv-kglobal-attnmlp-nov-20260815: 4.28x, loss 3.2269, Aug 16 mhep-abl6711-d768-sconv-klocal-stat-attnmlp-nov-20260815: 4.31x, loss 3.2261, Aug 16 mhep-abl6711-d768-experts-128-20260814-r6: 4.13x, loss 3.2360, Aug 16 mhep-abl6711-d768-input-output-skip-20260815: 4.50x, loss 3.2170, Aug 16 mhep-abl6711-d768-input-output-skip-mlp-20260815: 4.53x, loss 3.2159, Aug 16 mhep-abl6711-d768-input-mid-output-skip-mlp-20260815: 4.70x, loss 3.2117, Aug 16 mhep-abl6711-d768-input-mid-output-skip-mlp-adam-20260815: 5.00x, loss 3.2019, Aug 16 mhep-abl6711-d768-input-mid-output-skip-mlp-adam-0p1-20260815: 4.75x, loss 3.2114, Aug 16 mhep-abl6711-d768-input-mid-output-skip-mlp-adam-0p1-gelu-20260815: 4.83x, loss 3.2072, Aug 16 mhep-prpool-d768-send115-recv115-nov-polar8-20260815: 2.55x, loss 3.2790, Aug 16 mhep-prpool-d768-send115-recv115-nov-20260815-r2: 2.61x, loss 3.2772, Aug 16 mhep-prpool-d768-send115-recv115-nov-propmix-20260815: 3.41x, loss 3.2357, Aug 16 mhep-prpool-d768-send115-recv115-nov-polar8-safe101-20260815: 2.67x, loss 3.2723, Aug 16 mhep-abl6711-d768-adamh-outproj-grow-20260816: 4.31x, loss 3.2218, Aug 16 mhep-abl6711-d768-zloss-4x-20260816: 4.51x, loss 3.2175, Aug 16 mhep-abl6711-d768-outproj-adamw-20260816: 4.42x, loss 3.2200, Aug 16 mhep-abl6711-d768-outproj-adam-20260816: 4.12x, loss 3.2270, Aug 16 mhep-abl6711-d768-outproj-adamw-lrcoupled-20260816: 4.32x, loss 3.2239, Aug 16 mhep-abl6711-d768-outproj-cwd-20260816-r2: 4.34x, loss 3.2221, Aug 17 rav-ladder-d768-v2: 2.01x, loss 3.3469, Aug 18 rav-ladder-d768-v2-old-mix: 2.22x, loss 3.3341, Aug 18 MOE-PSC-102-d768-v5p8-gate1: 0.88x, loss 3.2663, Aug 19 rav-ladder-d768-v2-semantic-pi-56rows-candidate1-pergpu-prefetch32-w512-ts125g: 1.89x, loss 3.3628, Aug 19 rav-ladder-d768-v2-epsilon0-monotonic-pi-57rows-candidate1-pergpu-prefetch32-w512-ts125g: 2.40x, loss 3.3260, Aug 19 rav-ladder-d768-v2-epsilon0-pi-57rows-candidate1-pergpu-prefetch32-w512-ts125g: 2.80x, loss 3.3050, Aug 19 MOE-PSC-CTRL-102-d768-v5p8: 0.77x, loss 3.2950, Aug 20 l2-abl-d768-ragged: 6.22x, loss 3.1616, Aug 22 l2-abl-d768-ep: 2.77x, loss 3.3197, Aug 22 Aug 11 Aug 22
Best
10.26× harrier-nemotron-proportional-d768-full-krr-20260814 loss 3.1007
This week
6.22× l2-abl-d768-ragged loss 3.1616
Baseline
moe-v16-compute-opt-d768-1.70e+18
d1024 / 9.00e18 FLOPs
100 completed runs; 1 this period
baseline loss 3.1605
1x gb200-d1024-gqa-16h: 3.50x, loss 3.1378, Jul 19 gb200-d1024-mla-qlora0-v2: 2.96x, loss 3.1727, Jul 19 gb200-d1024-mla-12h: 2.95x, loss 3.1520, Jul 19 gb200-d1024-mla-init2-uq: 3.52x, loss 3.1540, Jul 19 gb200-d1024-mla-lrdrop-uq: 3.29x, loss 3.1587, Jul 19 gb200-d1024-mla-lrdrop-uk: 3.25x, loss 3.1607, Jul 19 gb200-d1024-mla-lrdrop-dq: 3.27x, loss 3.1598, Jul 19 gb200-d1024-mla-lrdrop-kr: 3.22x, loss 3.1588, Jul 19 gb200-d1024-mla-lrdrop-uv: 3.41x, loss 3.1554, Jul 19 gb200-d1024-mla-init2-dq: 3.29x, loss 3.1599, Jul 19 gb200-d1024-mla-lrdrop-dkv: 3.34x, loss 3.1560, Jul 19 gb200-d1024-mla-init2-uv: 3.21x, loss 3.1631, Jul 19 gb200-d1024-mla-init2-kr: 3.27x, loss 3.1558, Jul 19 gb200-d1024-mla-init2-uk: 3.32x, loss 3.1576, Jul 19 gb200-d1024-mla-init2-dkv: 3.26x, loss 3.1604, Jul 19 gb200-d1024-mla-16h: 2.64x, loss 3.1496, Jul 19 gb200-d1024-gqa2mla-step2: 2.15x, loss 3.2407, Jul 20 gb200-d1024-gqa2mla-step1: 3.29x, loss 3.1754, Jul 20 gb200-d1024-gqa2mla-step6: 1.88x, loss 3.2507, Jul 20 gb200-d1024-gqa2mla-step0: 3.74x, loss 3.1555, Jul 20 gb200-d1024-gqa2mla-step3: 1.95x, loss 3.2525, Jul 20 gb200-d1024-gqa2mla-step4: 1.75x, loss 3.2620, Jul 20 gb200-d1024-gqa2mla-step5: 1.79x, loss 3.2597, Jul 20 gb200-d1024-gqa2mla-step7: 1.52x, loss 3.1615, Jul 20 gb200-d1024-gqa2mla-step8: 3.09x, loss 3.1610, Jul 20 gb200-d1024-mla-kvslice-alt: 3.28x, loss 3.1601, Jul 20 gb200-d1024-mla-kvfreeze: 3.15x, loss 3.1650, Jul 20 gb200-d1024-mla-kvslice-first: 3.27x, loss 3.1594, Jul 20 gb200-d1024-mla-kvfreeze-ortho: 3.28x, loss 3.1605, Jul 20 gb200-d1024-rope-local512-g6-v1: 3.93x, loss 3.1597, Jul 20 gb200-d1024-relpos-local512-g6-v1: 1.04x, loss 3.1477, Jul 20 gb200-d1024-ropeall-local512-g6: 3.91x, loss 3.1607, Jul 20 gb200-d1024-ropeall-local1024-g6: 3.90x, loss 3.1573, Jul 20 h100-d1024-12L-conv-baseline-v3: 2.16x, loss 3.1994, Jul 24 h100-d1024-12L-conv-k-only-v3: 3.49x, loss 3.1963, Jul 24 h100-d1024-12L-conv-v-only-v3: 2.94x, loss 3.2056, Jul 24 h100-d1024-12L-conv-k-global-v3: 3.41x, loss 3.1960, Jul 24 h100-d1024-12L-conv-attn-only-v3: 3.27x, loss 3.2017, Jul 24 h100-d1024-12L-conv-all-k2-v3: 3.43x, loss 3.1931, Jul 24 h100-d1024-12L-conv-mlp-only-v3: 2.70x, loss 3.1997, Jul 24 h100-d1024-12L-conv-all-k3-v3: 2.85x, loss 3.1867, Jul 24 h100-d1024-12L-conv-all-k4-v3: 3.18x, loss 3.1858, Jul 24 h100-d1024-12L-conv-all-global-v3: 3.37x, loss 3.1922, Jul 24 h100-d1024-12L-pko-nope-v1: 3.44x, loss 3.1953, Jul 24 h100-d1024-12L-prope-v1: 3.23x, loss 3.2042, Jul 24 h100-d1024-12L-pko-prope-v1: 3.25x, loss 3.2027, Jul 24 h100-d1024-12L-kglobal-identinit-prope-v1: 3.47x, loss 3.1953, Jul 24 h100-d1024-12L-kglobal-pkoinit-prope-v1: 2.95x, loss 3.1987, Jul 24 h100-d1024-12L-kglobal-pkoinit-k4-prope-v1: 3.32x, loss 3.1995, Jul 24 h100-d1024-12L-base-sw2k-datakit-v1: 3.57x, loss 3.1850, Jul 24 h100-d1024-12L-pko-sw2k-datakit-v1: 3.71x, loss 3.1765, Jul 24 h100-d1024-11L-pko-11L-e256-datakit-v1: 3.69x, loss 3.0656, Jul 24 h100-d1024-11L-pko-11L-g4-e256-datakit-v1: 4.13x, loss 3.0544, Jul 24 h100-d1024-11L-g4-conv-k-only-e256-datakit-v1: 4.10x, loss 3.0542, Jul 24 h100-d1024-11L-g4-conv-k-global-e256-datakit-v1: 4.14x, loss 3.0531, Jul 24 h100-d1024-11L-g4-conv-attn-only-e256-datakit-v1: 3.99x, loss 3.0571, Jul 24 h100-d1024-11L-g4-conv-mlp-only-e256-datakit-v1: 4.17x, loss 3.0506, Jul 24 h100-d1024-11L-g4-conv-all-k3-e256-datakit-v1: 4.46x, loss 3.0387, Jul 24 h100-d1024-11L-g4-conv-all-global-e256-datakit-v1: 4.16x, loss 3.0482, Jul 24 h100-d1024-11L-g4-conv-all-k4-e256-datakit-v1: 4.47x, loss 3.0378, Jul 24 aug-hero-d1024-30x-lr1.4-v2: 3.06x, loss 3.1663, Aug 1 aug-hero-d1024-30x-lr0.7-v2: 3.05x, loss 3.1657, Aug 1 aug-hero-d1024-30x-lr1.2-v2: 3.19x, loss 3.1602, Aug 1 aug-hero-d1024-30x-lr1-v2: 3.37x, loss 3.1532, Aug 1 aug-hero-d1024-60x-lr0.7-v2: 3.37x, loss 3.0552, Aug 1 aug-hero-d1024-150x-lr0.7-v2: 3.05x, loss 2.9483, Aug 2 aug-hero-d1024-150x-lr0.85-v2: 3.14x, loss 2.9421, Aug 2 aug-hero-d1024-60x-lr0.85-v2: 3.56x, loss 3.0480, Aug 2 aug-hero-d1024-150x-lr1-v2: 3.28x, loss 2.9378, Aug 2 aug-hero-d1024-60x-lr1.4-v2: 3.39x, loss 3.0530, Aug 2 aug-hero-d1024-60x-lr1-v2: 3.57x, loss 3.0463, Aug 2 aug-hero-d1024-60x-lr1.2-v2: 3.50x, loss 3.0494, Aug 2 aug-hero-d1024-150x-lr1.4-v2: 3.12x, loss 2.9454, Aug 2 aug-hero-d1024-150x-lr1.2-v2: 3.29x, loss 2.9386, Aug 2 aug-hero-d1024-300x-lr0.85-v2: 2.72x, loss 2.8757, Aug 2 aug-hero-d1024-300x-lr1.4-v2: 2.79x, loss 2.8720, Aug 2 aug-hero-d1024-300x-lr1-v2: 2.82x, loss 2.8706, Aug 2 aug-hero-d1024-300x-lr0.7-v2: 2.52x, loss 2.8834, Aug 2 aug-hero-d1024-300x-lr1.2-v2: 2.88x, loss 2.8684, Aug 2 aug-hero-d1024-600x-lr0.85-v2: 2.23x, loss 2.8187, Aug 3 aug-hero-d1024-600x-lr1.2-v2: 2.33x, loss 2.8131, Aug 3 aug-hero-d1024-600x-lr1-v2: 2.25x, loss 2.8158, Aug 3 aug-hero-d1024-600x-lr1.4-v2: 2.34x, loss 2.8132, Aug 3 aug-hero-d1024-600x-lr0.7-v2: 2.03x, loss 2.8281, Aug 3 iso-3e18-d1024: 2.74x, loss 3.2236, Aug 6 iso-1e20-d1024: 7.22x, loss 2.8069, Aug 6 grug_xem_d1024_smoke_core_groups_two_anchor_unscaled: 0.04x, loss 4.3356, Aug 8 grug_xem_d1024_smoke_baseline: 0.04x, loss 4.3132, Aug 8 mhep-ladder-hist-20260808c-fsdp-chunk1-d1024: 17.40x, loss 2.8147, Aug 9 mhep-ladder-hist-20260808c-fsdp-chunk4-d1024: 12.80x, loss 2.8439, Aug 9 grug_xem_d1024_full_core_groups_two_anchor_unscaled: 0.90x, loss 3.0699, Aug 9 mhep-ladder-hist-noinit-20260808c-ep64-d1024: 4.32x, loss 2.9849, Aug 9 grug_xem_d1024_full_baseline: 1.10x, loss 3.0392, Aug 9 mhep-d1024-i768-k6-cf1p45-20260809: 5.28x, loss 2.9960, Aug 9 ppg-d1024-ep-cf133-b: 9.44x, loss 3.3850, Aug 12 ppg-d1024-ep-cf133-ctl4: 9.37x, loss 3.3873, Aug 12 mhep-abl-d1024-baseline-20260812: 5.40x, loss 2.9730, Aug 13 mhep-abl-d1024-e384-t8-half-smatch-20260812: 4.52x, loss 2.9709, Aug 13 mhep-ep8062-d1024-pool-w3-cf133-halfexp-20260814: 5.99x, loss 2.9573, Aug 16 rav-ladder-d1024: 4.03x, loss 3.0999, Aug 19 Jul 19 Aug 19
Best
17.40× mhep-ladder-hist-20260808c-fsdp-chunk1-d1024 loss 2.8147
This week
4.03× rav-ladder-d1024 loss 3.0999
Baseline
moe-v16-compute-opt-d1024-9.00e+18
d1280 / 2.83e19 FLOPs
85 completed runs
baseline loss 3.0065
1x muonh-matrix-baseline-adam-mask-d1280-2.83e19: 0.83x, loss 2.9888, May 11 muonh-nowarmup-d1280-2.83e19: 0.95x, loss 2.9706, May 13 muonh-gn-adamh-v1-d1280-2.83e19: 0.85x, loss 2.9855, May 15 muonh-may-recipe-lr-v1-d1280-R4-lr1p6: 1.51x, loss 3.4269, May 21 muonh-may-recipe-lr-v1-d1280-R4-lr0p4: 0.45x, loss 3.6450, May 21 muonh-may-recipe-lr-v1-d1280-R4-lr1p3: 1.59x, loss 3.4167, May 21 muonh-may-recipe-lr-v1-d1280-R4-lr0p7: 1.20x, loss 3.4664, May 22 muonh-may-recipe-lr-v1-d1280-R4-lr1p0: 1.59x, loss 3.4172, May 22 muonh-may-recipe-lr-v1-d1280-R20-lr0p4: 2.33x, loss 3.1066, May 22 muonh-may-recipe-lr-v1-d1280-R20-lr1p3: 3.44x, loss 3.0522, May 22 muonh-may-recipe-lr-v1-d1280-R20-lr1p0: 3.63x, loss 3.0448, May 22 muonh-may-recipe-lr-v1-d1280-R20-lr0p7: 3.42x, loss 3.0532, May 22 context-norm-no-xsa-gate2-v1-d1280-2.83e19: 0.76x, loss 3.0107, May 22 muonh-may-recipe-lr-v1-d1280-R20-lr1p6: 0.03x, loss 3.0692, May 22 grug_moe_mix_v4_path_r1_t050_d1280-2.83e+19: 0.90x, loss 2.9884, May 22 grug_moe_mix_v4_path_r1_t075_d1280-2.83e+19: 0.84x, loss 2.9962, May 22 muonh-may-recipe-lr-v1-d1280-R60-lr1p0: 4.20x, loss 2.8851, May 22 muonh-may-recipe-lr-v1-d1280-R60-lr1p6: 3.47x, loss 2.9081, May 22 muonh-may-recipe-lr-v1-d1280-R60-lr0p7: 4.18x, loss 2.8856, May 22 muonh-may-recipe-lr-v1-d1280-R60-lr0p4: 3.13x, loss 2.9211, May 22 muonh-may-recipe-lr-v1-d1280-R60-lr1p3: 3.87x, loss 2.8948, May 22 grug_moe_mix_v4_path_r1_t025_d1280-2.83e+19: 0.92x, loss 2.9851, May 24 muonh-may-recipe-lr-v1-d1280-R120-lr0p4: 3.24x, loss 2.8338, May 25 muonh-may-recipe-lr-v1-d1280-R120-lr0p7: 4.12x, loss 2.8063, May 25 muonh-may-recipe-lr-v1-d1280-R120-lr1p6: 3.32x, loss 2.8307, May 25 muonh-may-recipe-lr-v1-d1280-R120-lr1p3: 3.76x, loss 2.8162, May 25 muonh-may-recipe-lr-v1-d1280-R120-lr1p0: 3.99x, loss 2.8097, May 25 grug-moe-isoflop-v3e18-d1280-v1: 1.34x, loss 3.2983, May 27 grug-moe-isoflop-v3e19-d1280-v1: 3.27x, loss 2.9045, May 29 marin-big-run-moe_may_compute_opt_d1280: 2.04x, loss 2.8963, Jun 3 moe_may_compute_opt_d1280_ep1: 1.99x, loss 2.8857, Jun 5 moe_may_compute_opt_d1280_ep2_16kctx_long_yarn_mscale01_from13k: 1.51x, loss 2.8572, Jun 5 moe_may_compute_opt_d1280_ep1_16kctx_long_yarn_mscale01_from13k: 1.47x, loss 2.8473, Jun 5 moe_may_compute_opt_d1280_ep1_longmino_from13k: 1.05x, loss 2.9659, Jun 5 moe_may_compute_opt_d1280_ep1_longmino_halfmix_from13k: 1.94x, loss 2.8887, Jun 5 moe_may_compute_opt_d1280_ep2_longmino_from13k: 1.08x, loss 2.9776, Jun 5 moe_may_compute_opt_d1280_ep2_longmino_halfmix_from13k: 2.03x, loss 2.8979, Jun 5 moe_may_compute_opt_d1280_ep8_longmino_from13k: 0.71x, loss 3.0018, Jun 5 moe_may_compute_opt_d1280_ep8_longmino_halfmix_from13k: 1.32x, loss 2.9211, Jun 5 moe_may_compute_opt_d1280_ep8_32kctx_long_yarn_mscale01_halfmix_from13k: 0.66x, loss 2.8675, Jun 5 moe_may_compute_opt_d1280_ep1_seq8k: 1.82x, loss 2.8664, Jun 8 mtp-d1280-baseline: 3.76x, loss 2.9397, Jul 15 mtp-d1280-densestep: 2.46x, loss 2.9306, Jul 15 mtp-d1280-step: 2.42x, loss 2.9278, Jul 15 mtp-d1280-linear: 2.44x, loss 2.9270, Jul 15 aug-d1280-lin-lr1p1: 6.09x, loss 2.9130, Jul 29 aug-d1280-lin-lr0p9: 6.06x, loss 2.9135, Jul 29 aug-d1280-lin-lr1p0: 6.09x, loss 2.9124, Jul 29 aug-d1280-1sqrt-lr1p3: 6.07x, loss 2.9136, Jul 29 aug-d1280-1sqrt-lr1p4: 5.94x, loss 2.9135, Jul 29 aug-d1280-1sqrt-lr1p5: 6.06x, loss 2.9130, Jul 29 aug-hero-d1280-30x-lr1.2-v2: 7.66x, loss 2.9963, Aug 1 aug-hero-d1280-30x-lr0.7-v2: 7.02x, loss 3.0059, Aug 1 aug-hero-d1280-30x-lr1-v2: 7.71x, loss 2.9939, Aug 1 aug-hero-d1280-30x-lr0.85-v2: 7.54x, loss 2.9967, Aug 1 aug-hero-d1280-30x-lr1.4-v2: 7.05x, loss 3.0055, Aug 1 aug-hero-d1280-60x-lr1.4-v2: 7.80x, loss 2.9062, Aug 1 aug-hero-d1280-60x-lr1.2-v2: 8.18x, loss 2.9002, Aug 1 aug-hero-d1280-150x-lr1.4-v2: 7.15x, loss 2.8068, Aug 2 aug-hero-d1280-150x-lr1.2-v2: 7.54x, loss 2.8017, Aug 2 aug-hero-d1280-150x-lr0.7-v2: 6.80x, loss 2.8129, Aug 2 aug-hero-d1280-150x-lr0.85-v2: 7.16x, loss 2.8065, Aug 2 aug-hero-d1280-150x-lr1-v2: 7.48x, loss 2.8025, Aug 2 aug-hero-d1280-60x-lr1-v2: 8.32x, loss 2.8993, Aug 2 aug-hero-d1280-60x-lr0.85-v2: 8.06x, loss 2.9028, Aug 3 aug-hero-d1280-60x-lr0.7-v2: 7.56x, loss 2.9100, Aug 3 aug-hero-d1280-300x-lr0.7-v2: 5.89x, loss 2.7546, Aug 3 aug-hero-d1280-300x-lr1-v2: 6.67x, loss 2.7415, Aug 4 aug-hero-d1280-300x-lr1.4-v2: 6.44x, loss 2.7437, Aug 4 aug-hero-d1280-300x-lr1.2-v2: 6.66x, loss 2.7399, Aug 4 aug-hero-d1280-300x-lr0.85-v2: 6.24x, loss 2.7464, Aug 4 aug-hero-d1280-600x-lr1.4-v2: 5.25x, loss 2.6917, Aug 5 aug-hero-d1280-600x-lr0.85-v2: 4.92x, loss 2.6979, Aug 5 aug-hero-d1280-600x-lr0.7-v2: 4.62x, loss 2.7052, Aug 6 iso-3e18-d1280: 2.15x, loss 3.2895, Aug 6 aug-hero-d1280-600x-lr1.2-v2: 5.23x, loss 2.6918, Aug 6 aug-hero-d1280-600x-lr1-v2: 5.20x, loss 2.6924, Aug 6 iso-3e19-d1280: 0.00x, loss 11.7618, Aug 6 iso-1e20-d1280: 11.82x, loss 2.7716, Aug 6 abl-ec-d1280-c4: 3.24x, loss 2.9120, Aug 7 qb-bias-d1280: 3.59x, loss 2.8980, Aug 7 abl-ec-d1280-c1: 3.63x, loss 2.8979, Aug 8 abl-ec-d1280-c2: 3.23x, loss 2.9061, Aug 8 grug_xem_d1280_full_core_groups_two_anchor_unscaled: 1.05x, loss 2.9297, Aug 11 grug_xem_d1280_full_baseline: 1.31x, loss 2.8999, Aug 11 May 11 Aug 11
Best
11.82× iso-1e20-d1280 loss 2.7716
This week
no completed point
Baseline
moe-v16-compute-opt-d1280-2.83e+19

Top 15 runs (by FLOPs) this week (completed, running, crashed)


The week’s 1.22M chip-hours split between two live hero runs and a burst of GB200 infrastructure work. The 67B-A2B 10T hero run continues on 1,024 TPU v4 chips at 951K chip-hours, having processed 8.96T of 10.07T tokens (89%). Paloma macro loss holds at 2.222, comfortably below the 2.269 preregistered target. The third cooldown branched at step 141k finished this week at Paloma macro 2.177. A new context-extension run branched at step 156k started August 22, extending the sequence length from 8K to 262K tokens; it is running and already shows Paloma macro 2.210.

The main event is the new 535B-A23B hero run on 704 NVIDIA GB200 GPUs, launched August 20 by Rafal Wojdyla after a crashed predecessor the day before. At 56K chip-hours it has processed 480B tokens with 21.2% MFU (Model FLOPs Utilization) and Paloma macro loss of 2.626, still early in its training curve. Over 20 merged PRs this week supported the run’s infrastructure: fused cross-entropy bf16 backward, fused ShortConv (short convolution), and hoisted router all-reduces in #8385 (Larry Dial); concurrent checkpoint array reads in #8505 and replicated-shard deduplication in #8589 that cut a 12.85x checkpoint read amplification down to 1x (Russell Power); a bounded post-restore barrier in #8538 after one stalled rank held 175 others for 4.5 hours (Matt Wittmann); and jemalloc as the default allocator in #8550 to address recurrent checkpoint-save OOMs #8506. The host-memory OOM during checkpointing remains under active investigation in #8599.

Matt Wittmann’s #8549 draft introduces a ragged all-to-all expert-parallel (EP) MoE backend as a replacement for the pooled-wave transport. An A/B at hero shape restored from the live run’s step-6000 checkpoint shows matching throughput (22.5 vs 22.5 MFU), 180x lower token drop rate (0.015% vs 2.67%), and 10 GiB lower peak device memory. In the Agent MoE scaling ladder, the d768 ragged ablation completed at 6.22x effective speedup, the highest among the 13 runs finished this period. Will Held updated the Harrier data mix with monotonic PI (proportional importance) weights in #8452, and the first d1024 ladder rung (rav-ladder-d1024) completed at 4.03x speedup. David Hall’s restore-slots soak test validated checkpoint restore correctness for the hero shape, converging train loss to near zero.

Run User Hardware(?) Hours(?) FLOP Budget(?) Loss BPB(?)
#6705 moe_67b_a2b_d2560_ep1_rep8_bs8192_seq8192_sw2k_v4_2048_muon_resume15k_v2_10T Larry Dial TPU v4
(1024 chips)
38.7d 1.83e23 model
9.97e23 HW (18%)
BPB: 0.656
#6705 moe_67b_a2b_d2560_ep1_rep1_ctx4_bs256_seq262144_ctxext_step156k Larry Dial TPU v4
(1024 chips)
20.2h 1.40e23 model
7.96e23 HW (18%)
BPB: 0.634
#6689 hero-12d8b6f0-dee637 Rafal Wojdyla NVIDIA GB200
(704 chips)
3.3d 7.18e22 model
3.39e23 HW (21%)
BPB: 0.790
#6705 moe_67b_a2b_d2560_ep1_rep8_bs1024_seq65536_sw2k_v4_2048_muon_cooldown_step141k Larry Dial TPU v4
(1024 chips)
6.5d 5.59e22 model
2.96e23 HW (19%)
BPB: 0.636
#6689 mhep-d2560-L24-pool-cf115-8of384-4rack-20260816-r2 Larry Dial NVIDIA GB200
(256 chips)
1.3d 9.68e21 model
6.16e22 HW (16%)
BPB: 0.848
rav-ladder-d2048-v3 Rafal Wojdyla NVIDIA GB200
(704 chips)
12.2h 7.50e21 model
5.46e22 HW (14%)
BPB: 0.843
mhep-lmhead-adamh-d1536-3rack-20260816 Larry Dial NVIDIA GB200
(192 chips)
12.6h 1.74e21 model
1.62e22 HW (11%)
BPB: 0.844
mhep-lmhead-cwd-d1536-3rack-20260816-r4 Larry Dial NVIDIA GB200
(192 chips)
12.8h 1.74e21 model
1.61e22 HW (11%)
BPB: 0.849
rav-ladder-d1536 Rafal Wojdyla NVIDIA GB200
(384 chips)
8.1h 1.83e21 model
1.57e22 HW (12%)
BPB: 0.888
#8508 qb8033-5r-hist10k-p Matt Wittmann NVIDIA GB200
(320 chips)
5.4h 3.14e21 model
1.52e22 HW (21%)
#8508 qb8033-5r-topk-p2 Matt Wittmann NVIDIA GB200
(320 chips)
5.4h 3.14e21 model
1.51e22 HW (21%)
#8508 hero-restore-slots-8508-soak-20260821-1909 David Leo Wright Hall NVIDIA GB200
(64 chips)
1.6h 3.03e21 model
1.39e22 HW (22%)
#8508 hero-restore-slots-8508-smoke10-20260821-1850 David Leo Wright Hall NVIDIA GB200
(64 chips)
0.2h 2.85e21 model
1.31e22 HW (22%)
#6689 hero-20260819 Rafal Wojdyla NVIDIA GB200
(704 chips)
2.2h 2.67e21 model
1.25e22 HW (21%)
rav-ladder-d1536-v2-semantic-pi-56rows-candidate1-pergpu-prefetch32-w512-ts125g Will Held NVIDIA GB200
(384 chips)
5.4h 1.34e21 model
1.15e22 HW (12%)
BPB: 0.914
Merged PR Open PR Draft PR Closed PR Open issue Closed issue

Keyboard shortcuts

?
Toggle this help
j / k
Next / previous section
t
Toggle details in current section
s
Cycle sort order in current section
o
Open current epic on GitHub
m
Open current milestone on GitHub
M
Open milestones list on GitHub
Data: weekly-data-2026-08-17_2026-08-23.json · sections-2026-08-17_2026-08-23.json · wandb-flops-2026-08-17_2026-08-23.json · tpu-usage-2026-08-17_2026-08-23.json · token-counts-2026-08-17_2026-08-23.json · cluster-status-2026-08-17_2026-08-23.json · discord-2026-08-17_2026-08-23.json · agent-moe-2026-08-17_2026-08-23.json