Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

What is the best way in terms of price/convenience ratio to run the 70B model on the cloud? Are there any providers offering out-of-the box setups?


I think using this project https://github.com/ggerganov/llama.cppav

on a CPU machine with AVX instructions would be a better bang for your buck than GPU. Depends on if your use case can tolerate the latency




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: