Optimal packings of 15-20 equal circles in a circle
A short note on proving optimal packings of 15, 16, 17, 18 and 20 circles in a circle: history, how the proofs were found with an LLM, and the main ideas.
A short note on proving optimal packings of 15, 16, 17, 18 and 20 circles in a circle: history, how the proofs were found with an LLM, and the main ideas.
A discussion of MoE load balancing, understanding quantile balancing, and generalizing and improving upon it. We show two improvements - one for improving load balancing/steering, and another for improving loss by making load balancing strength configurable.
An elementary finite-time account of why wide networks trained on the same minibatches develop matching local loss fluctuations, and how width and batch size control initialization, data, and interaction noise, with implications on scaling.
A short note on Adam's update size, the coupling of beta1 and beta2, bias correction, epsilon, and how these relate to stability and warmup.
A refinement of Sawin's explicit unit-distance lower bound, mostly by GPT 5.5 Pro, for the Erdős problem recently solved by OpenAI's internal model.
A short note on why RMS-matching Muon to AdamW can break width transfer, causing either undertraining or instability depending on scale.
A research note on what breaks in long-context attention, deriving a logit scaling, with QK-norm, hybrid/local attention, gating, and small-scale experiments.
An intuition-building tour of dynamical systems, using canonical examples to connect feedback, thresholds, coupling, noise, reinforcement, spatial structure, and phase transitions.
A writeup of porting the modded nanoGPT speedrun to pure JAX on TPU v6e, including hardware bottlenecks, bugs, optimizations, and open performance questions.
A theoretical comparison of normalized gradient descent, Muon, and Adam-style updates on a tractable matrix optimization toy problem, showing finite-time convergence and why Muon's guarantees come out nicer than Adam's.