What Losing Responses Have in Common
Human preference votes contain enough structure to calibrate a quality judge from scratch. Applied to the losing responses, that judge maps …
Read More
Testing gear on mountains. Testing ideas with mountains of data.
Gear reviews and data projects from Boulder, Colorado
Human preference votes contain enough structure to calibrate a quality judge from scratch. Applied to the losing responses, that judge maps …
Read More
A light, agile trail racer with excellent dry-surface grip and a secure heel, best suited to short-to-mid-distance mountain efforts.
Read Full ReviewDetecting that an AI system is failing is the easier problem. A failure taxonomy maps observable loss signals to the system layer …
Read MoreWhen GPT-4 launched in March 2023, the topic mix of real user prompts shifted in the exact direction the model quality gaps would predict: …
Read MoreCalifornia Denti-Cal records show Anaheim's payment intensity rose after the March 2015 ownership transition while every peer office fell or …
Read MoreThe model quality framework originally relied on scalar ratings and active days. By replacing those inputs with ARC trajectory scores and …
Read More