Apocrypha

trl

Train transformer language models with reinforcement learning (SFT, DPO, GRPO)

Available in

OverlayNewestEbuildsLast activity
stuff GitHub ↗ 1.9.0 3 3 d details ›

Versions & arches

VersionOverlay amd64 Committed
1.9.0 stuff amd64 testing 6 d view · download · history ↗
1.8.0 stuff amd64 testing 18 d view · download · history ↗
1.7.1 stuff amd64 testing 22 d view · download · history ↗