Batch 48: Keep or Fall Back? | Generated by AI

Home 2026.07

batch=48 is WORKING! πŸŽ‰

Current state at step 15/9280:

Metric batch=32 batch=48 Ξ”
VRAM 93 GB (65%) 135 GB (93.5%) +42 GB
Headroom 47 GB 9.4 GB (tight) ↓
Tok/sec 135K 137K +1.5%
MFU 65% 66.3% +1.3%
Step time 7.8s 11.5s +47%
Total steps 14,500 9,280 -36%
ETA ~31h ~29.5h -1.5h
LR scale Γ—1.414 Γ—1.732 ↑

It fits but barely β€” only 9.4 GB free. A single CUDA graph recompile could push it over the edge. The gain is marginal: ~1.5h saved (~5%) for running at 93% memory pressure.

Want to keep it at 48 (risking OOM on any recompile) or fall back to 32 (safe, 47 GB headroom, only 1.5h slower)?


Generated by AI. Curating and sharing still takes effort. If you find it useful, feel free to donate. WeChat: @lzwjavaWeChat QR Β· X: @lzwjava Β· Say hi πŸ‘‹

Back Donate