Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engine

(modal.com)

1 points | by gmays 6 hours ago ago

No comments yet.