Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

That's a 2.4T model, how would they reduce this to 35B and still give some accuracy? That's a completely different arch.


there's been a lot of research about reducing models by taking out layers; there's also using it to train smaller models by optimizing parameters.

I dont see most model building as anything more than a pig at a slop troth, despite the level of sophistication; they're still rarely pruning the input beyond random sampling.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: