Non-HF data sources: github.com/suttacentral/bilara-data, github.com/cbeta-org/xml-p5, github.com/OpenITI/RELEASE Also a private dataset based on personal Discord logs for indefinite context dialogue training.
Long documents with varying forms of long-range dependencies for the objective, replay data, and a dash of extra math and code thrown in.
Rank 128 LoRA adapter for long context training, 50.0M tokens, 180 steps, whole-document training with 32k gradient windows (docs ≤256k).
5e-5 learning rate, WSD schedule.
Flash attention path utilized, consistency with the standard modeling paths cross-checked.
Targets and removes pollution penalty to prediction from long state, tested at 56k/196KB. May also help with handling token-hungry languages and less continuous languages.
Slight tax on English harness benchmarks; the x0.7 adapter has a lower tax but does not address state pollution as effectively.
Model tree for Lambent/RWKV7-1.5B-midtrain50-docs-lora
Base model
RWKV/RWKV7-1.5B-20260805