SIGN IN SIGN UP

feat: Router.preload so routing never pays a model load (v0.3.2)

A cold checkpoint build costs seconds; language detection costs microseconds.
At the default max_loaded=1, traffic that alternates languages rebuilt a model
on every request -- which is what made routing feel slow in the demo Space.

Router(preload=True) builds every checkpoint up front and raises max_loaded to
fit them, so the LRU cannot immediately evict what it just built.
router.preload([...]) does a subset.

Measured on CPU with the Space's own workload: preloaded requests run
193-464 ms with no model load on any request; without it each language switch
paid a 4-6 s rebuild.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
N
Nandakishor committed
27cbc8c8d29ebbc9f3d8cb0ca835491fd14e6396
Parent: c369320