feat: Router.preload so routing never pays a model load (v0.3.2)
A cold checkpoint build costs seconds; language detection costs microseconds. At the default max_loaded=1, traffic that alternates languages rebuilt a model on every request -- which is what made routing feel slow in the demo Space. Router(preload=True) builds every checkpoint up front and raises max_loaded to fit them, so the LRU cannot immediately evict what it just built. router.preload([...]) does a subset. Measured on CPU with the Space's own workload: preloaded requests run 193-464 ms with no model load on any request; without it each language switch paid a 4-6 s rebuild. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
N
Nandakishor committed
27cbc8c8d29ebbc9f3d8cb0ca835491fd14e6396
Parent: c369320