May 21st, 2024 [1/2]
In partnership with

🔬Core research improving LLMs!
Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
💡Why?: Current methods of training LLMs on fixed-length token sequences. This method leads to cross-document attention within a sequence, which is not an ideal learning signal and is computationally inefficient. Additiona…


