AI / Model series
The Hikaru Series.
The testbed for our central idea. Hikaru exists to find out whether a model trained with Explanatory Reinforcement Learning behaves measurably differently from one trained on a binary reward.
What's here?
Models in this series.
What the series is for
Before we build anything large, we want to know that our training method works. Hikaru is that proof. Each model in the series is a finetune of an existing base, onto which we apply our ERL training stage. That keeps the experiment focused on what ERL changes, rather than entangling the results with the cost and variance of pretraining.
Hikaru is deliberately small and deliberately honest about its scope. It is not the model we hope to ship to the world; it is the model that tells us whether the rest of the plan is sound.
Hikaru's pace slowed once we set a clear order of priority across projects; see Evaluating scope for why. Work continues, just more slowly than Deriva.
Related
- Hikaru-1, the model page
- How ERL works, the method Hikaru-1 is built to test
- Hikaru-1 enters development, the announcement
- Evaluating scope, why Hikaru slowed down
- All model series: Arcus, Deriva, and Nero