We're still just scratching the surface of how much better, faster, cheaper we can make inference for open-source models!

We’re still just scratching the surface of how much better, faster, cheaper we can make inference for open-source models!

So super excited to see companies like @zml_ai by @steeve come out of stealth with their inference engine integrated with @huggingface as the storage layer.

Let’s push open-source tokens to the max!

https://bender.layer3.press/articles/019f45c2-155e-05e7-70b1-05ca82888c3d

Write a comment