I've noticed a shift in how I work over the past year, and I'm not 100% sure it's a good one.
A year ago, I'd use AI to speed up execution like writing boilerplate code, drafting an email or summarizing a doc. The thinking was still mine. Now, more and more often, I catch myself asking AI to decide things too: which architecture to pick, how to frame a pricing page, even how to respond to a tricky customer email. The line between "helping me think" and "thinking for me" has gotten blurry.
Here's what worries me a bit: the moments where I learn the most are usually the moments where I struggle first. If AI removes the struggle, does it also remove the learning? Or is that just an old-fashioned way of looking at it - the same way people worried calculators would ruin math skills, and mostly they didn't?
Do you think AI has made you better at thinking, or just faster at producing? And any moment where relying on AI backfired because you skipped the "struggle" step?
Let me start from where i'm sitting: i talk to founders, hiring managers, and recruiters every week, and the same conversation keeps coming up. nobody trusts their pipeline anymor
Application volumes have doubled since 2021. Cheating on technical assessments doubled in just six months last year, from 15% to 35% of candidates. One in three hiring managers has personally caught a fake identity in an interview. @Anthropic had to rewrite their interview questions after so many people cheated during the interview with Claude.
I built a benchmark that runs nine WordPress backup plugins head to head on identical sites. One of the nine is mine, and it currently sits second. That is an obviously compromised position, so I have spent more time on what makes a vendor-run benchmark believable than on the measuring.
Three things I landed on, and I am not sure any of them are enough.
Publish the runs you threw away. 581 completed runs, 180 scored, 401 excluded, and every excluded one is browsable with its reason attached. Most are boring reruns rather than embarrassing failures, and saying so is part of the point: 360 superseded, 41 excluded for cause. The honest number includes the boring ones.
Say when a lead is not a lead. Mine sits five points above third place with plus or minus 113 error bars. That is a tie, and the page says tie rather than letting me quote the higher number.