AI  /  Model series

The Hikaru Series.

The testbed for our central idea. Hikaru exists to find out whether a model trained with Explanatory Reinforcement Learning behaves measurably differently from one trained on a binary reward.

Hikaru Series card

What the series is for

Before we build anything large, we want to know that our training method works. Hikaru is that proof. Each model in the series is a finetune of an existing base, onto which we apply our ERL training stage. That keeps the experiment focused on what ERL changes, rather than entangling the results with the cost and variance of pretraining.

Hikaru is deliberately small and deliberately honest about its scope. It is not the model we hope to ship to the world; it is the model that tells us whether the rest of the plan is sound.

Hikaru's pace slowed once we set a clear order of priority across projects; see Evaluating scope for why. Work continues, just more slowly than Deriva.

Related