Skip to content
#

training-dynamics

Here are 39 public repositories matching this topic...

Two-parameter Weibull lens on transformer weights (shape k, scale λ): 7-family benchmark, AdamW training dynamics (three-force λ evolution), a data-predictability law for λ growth, and a 76-run grid showing weight-scale growth tracks training effort, not learning quality. npm-weibull-py + database + code, arXiv:2605.18898/2606.19367/2608.23573

  • Updated Aug 26, 2026
  • Python

TMLR 2026 | Mechanistic interpretability: attention-head binding (EB*) as a marker of concept emergence. 7 models, 5 architectures (Pythia 160M–2.8B, OLMo-1B, CRFM GPT-2, SmolLM3-3B, Qwen2.5-1.5B), 41 terms.

  • Updated Jun 9, 2026
  • Python

Add this topic to your repo

To associate your repository with the training-dynamics topic, visit your repo's landing page and select "manage topics."

Learn more