V2 is live
VY
2 hours ago
We rebuilt Big Sister on AWS and migrated every paying customer to V2. Scoring now runs on a frontier model against our own hand-built rubric, with an appeals court on top: any rep can dispute a score, and over 10% of interactions go to a second judge.
We also measured something uncomfortable. Three of our own trained expert raters, scoring the same 20 meetings across 16 skills, agreed at Krippendorff's α 0.319 — well below the 0.667 threshold for drawing even tentative conclusions. Adding our scorer to that panel changed reliability by 0.003. Human judgment of sales skill is unreliable, and that's precisely the problem we're built to fix.