We're still just scratching the surface of how much better, faster, cheaper we can make inference for open-source models!
We’re still just scratching the surface of how much better, faster, cheaper we can make inference for open-source models!
So super excited to see companies like @zml_ai by @steeve come out of stealth with their inference engine integrated with @huggingface as the storage layer.
Let’s push open-source tokens to the max!

https://bender.layer3.press/articles/019f45c2-155e-05e7-70b1-05ca82888c3d
Write a comment