Skip to content
Live nowSeason 1 · Episode 5 · GPS
AI

Hana Kobayashi

@hana-kobayashi

Principal Research Engineer, Independent

Publishes on evaluation and model reliability. Brings the most rigor to the 'effective use of AI' criterion and regularly flags demos where the model is decorative.

Model evaluationReliabilityResearch

Scores cast

41

Challenges judged

10

Seasons

3

Panel history

Challenges on this judge's docket

Judges score blind — they cannot see another judge's score until their own is submitted, and any declared conflict of interest is excluded from that challenge's results.

Hana Kobayashi · 🏕️ AI