AI & ML interests

None defined yet.

Recent Activity

Raghav-SinghalĀ  updated a Space about 11 hours ago
dlab-spp/README
Raghav-SinghalĀ  updated a dataset about 11 hours ago
dlab-spp/corpus-verification
Raghav-SinghalĀ  updated a dataset about 11 hours ago
dlab-spp/corpus-1T-manifest
View all activity

Organization Card

SPP: annotate pretraining data with normative reflections, inject at different pretraining stages, evaluate alignment and safety

Alignment — and the assistant identity itself — is normally introduced only after pretraining, once behavioral priors are already set. SPP installs the desired persona from token zero instead: we define it through normative values in a constitution, generate first-person moral reflections grounded in that constitution, and insert them throughout the pretraining corpus behind an <assistant> token. Post-training then binds the chat assistant identity to the installed persona. Pretraining up to 3B on 500B tokens, SPP improves constitution following and jailbreak robustness while preserving capabilities — and when the data arrives matters: models trained with reflections from token zero prioritize values differently and take fewer risky actions in out-of-distribution moral dilemmas than models given the exact same data only at the end of pretraining, an advantage that grows with scale.

šŸ“„ Paper: Synthetic Persona Pretraining: Alignment from Token Zero

Collections

šŸ“¦ Pretraining Datasets — the reflection data, the corpus selection manifest, safety scores, and verification files.

šŸ¤– Models — 3B Ā· Models — 1.7B — all five recipes, at both scales. We release all pretraining checkpoints, base, and instruct models, at both scales

šŸ’¬ Post-training Dataset — SP-SFT, the mixture that performs persona binding.

šŸ“Š Evals — ConstitutionEval and an audited AIRiskDilemmas.

From EPFL DLAB.

Citation

@misc{minder2026syntheticpersonapretrainingalignment,
      title={Synthetic Persona Pretraining: Alignment from Token Zero},
      author={Julian Minder and Viktor Moskvoretskii and Raghav Singhal and Difan Jiao and Andy Arditi and Shaobo Cui and Yiderigun Borjigin and Kartik Bali and Stefan Krsteski and Harsh Raj and Huu Nguyen and Jannik Brinkmann and Ashton Anderson and Roland Aydin and Robert West},
      year={2026},
      eprint={2608.13482},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2608.13482},
}