Hacker News
new
|
past
|
comments
|
ask
|
show
|
jobs
|
submit
login
OkGoDoIt
on July 18, 2023
|
parent
|
context
|
favorite
| on:
Llama 2
What's the best way to run inference on the 70B model as an API? Most of the hosted APIs including HuggingFace seem to not work out of the box for models that large, and I'd rather not have to manage my own GPU server.
Guidelines
|
FAQ
|
Lists
|
API
|
Security
|
Legal
|
Apply to YC
|
Contact
Search: