Hacker Newsnew | past | comments | ask | show | jobs | submit | fromlogin
Verifying coding AIs for LLM powered software (aunhumano.com)
2 points by mfalcon 7 months ago | past
Evaluating Agents (aunhumano.com)
42 points by mfalcon on Sept 3, 2025 | past | 9 comments
Building an AI judge for classification tasks (aunhumano.com)
2 points by mfalcon on May 2, 2025 | past
Building self improving negotiation agents (aunhumano.com)
1 point by mfalcon on March 17, 2025 | past
The time of evaluation driven development (aunhumano.com)
3 points by mfalcon on Feb 13, 2025 | past | 3 comments

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: