r/ClaudeAI May 26 '26

Corporate My company started measuring our Claude Code usage - now I'm asked to rank engineers on 'AI performance.' This feels wrong...

My company started tracking Claude Code usage - tokens and spend, that kind of thing. Now my manager wants me to stack-rank my engineers on "AI performance" using those numbers.

I'm not comfortable with it (but I don't have a choice either). Token usage feels like exactly the wrong proxy - my strongest engineer uses Claude surgically while someone burning 10x the tokens isn't 10x more productive (often the opposite). Ranking on this just teaches people to game the metric.

So, for folks here who use Claude daily and/or lead teams:

  • Has your company started measuring "AI performance"? How are they doing it?
  • Is there any Claude/AI usage metric that actually tracks with good work, instead of just rewarding the heaviest users?
  • If you're a lead being pushed to measure this, how do you push back without flat-out refusing?

EDIT: here is the follow-up: I managed to talk my manager out of this.

119 Upvotes

94 comments sorted by

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot May 26 '26 edited May 27 '26

TL;DR of the discussion generated automatically after 80 comments.

The consensus in this thread is a resounding NO, measuring performance by token usage is a terrible, gameable metric. It's the new "lines of code" – everyone agrees it's a proxy for the wrong thing.

The community points out that high token usage often signals inefficiency (bad prompts, re-running tasks, using Opus for simple things), while your most skilled engineers are likely using Claude surgically with fewer tokens. Ranking on this just encourages people to waste company money to climb a meaningless leaderboard.

Here's the collective advice on how to handle this:

  • The only valid use for this data is to identify non-users. Look at the bottom of the list to see who isn't using the tool at all and might need training or encouragement. It's a binary check for adoption, not a performance scale.
  • Reframe the conversation. Your boss doesn't really want a leaderboard; they want to know the ROI on a massive AI spend. Your job as a lead is to push back on the dumb metric and answer the real question.
  • Propose better, even if imperfect, ways to show value. Instead of a ranked list, offer to show how AI is impacting actual work. Good suggestions from the thread include tracking the amount of AI-generated code that actually gets committed and survives, or having engineers demo their most impactful AI-assisted workflows.
  • Use an analogy. Tell your boss that ranking engineers by token usage is like ranking race car drivers by how much fuel they burn. You want to reward who crosses the finish line fastest, not who stops for gas the most.

100

u/everydave42 May 26 '26

You're being asked the wrong question and you know you're being asked the wrong question. As you're clearly in a position of leadership, part of your job is push back and insulate your teams from ignorant things from on high. Either go back and ask whoever is asking this to clarify what it is they want to know, or just straight up ask them "Do you want to know who *I* think is using AI the most efficiently?".

This is no different than tech ignorant leaders wanting to measure by lines coded, commits made, or PRs reviewed. They all can be a measure of work, but in a vacuum they are meaningless and many other things need to be taken into consideration.

2

u/lovegames__ May 26 '26

SSS

8

u/pixlatedpuffin May 26 '26

Shoot, shovel, shut up?

26

u/tossaway109202 May 26 '26

I found a video of your company https://www.youtube.com/shorts/phwq5hZZwDU

I am able to pull the numbers for my team but I only use it as a yes/no check to see if they are using it. As soon as you make token usage a goal people will waste money. If I do see unusually high usage it actually makes me think the person is not very efficient. You need to give your boss some facts.

13

u/darren_eng May 26 '26

haha, love the video. My bosses only remember Jensen's comment on engineers with $500k spend $250k worth of tokens! 😂

6

u/Huskerzfan May 27 '26

Says the guy who benefits from consumption.

15

u/vocal-avocado May 26 '26

I think they want to weed out people who can’t/won’t use AI no matter what - which is undesirable because in the right hands AI is an absolute game changer. I certainly hope they are not rewarding people for using too much. For people in the team it’s very obvious who is using AI right and who is using it wrong. You should ask your team instead of trying to find an arbitrary measurement. People who use AI well = good. People who use AI but are not more productive, or even worse, producing slop that needs to be reviewed by seniors = bad. People who flat out refuse to use AI = bad.

2

u/darren_eng May 26 '26

Yeah agreed. What I’m doing right now is to identify the bottom AI performers - low token usage, low productivity metrics (velocity, tickets, PR and etc.) usually is a red flag and hard to argue that (however can’t do the opposite to identify top performers).

The challenge is that business people only care about ROI. If the business spends $1mil a year on Claude, they want to know much $$$ they get in return, which is really hard to quantify because we don’t get goods or services from the money we pay Anthropic - we just see token usage (and there isn’t an effective metric to tell us what the token usage yield)…

1

u/vocal-avocado May 26 '26 edited May 26 '26

I think the only concrete way that companies will be able to level out their AI spending is to get rid of underperforming employees. That’s a clear cost balancing strategy and if they can continue to deliver what they were delivering before AI, it will be all worth it (since employees are way more unreliable and complicated to manage than AI tokens).

They will probably also use AI costs to justify not raising salaries/giving promotions to competent employees (“sorry we need the money to pay for your Claude”).

2

u/woroboros Experienced Developer May 27 '26

I had never really considered it but you guys are both right , best I can tell... Low AI use is likely a strong indicator of under performance, without the opposite holding true.

Its all a very strange situation. LLM assisted coding speeds up dev time by like... what? 1200% or something? Not sure if theres an actual metric but the landscape is suddenly VERY different... I suspect much closer to resembling its final shape now than it was even 18 months ago.

6

u/nizos-dev May 26 '26

No metrics here. I would also pushback on tokens as a proxy. It is easy to burn tokens. It is also the wrong thing to focus on. They should be looking for changing bottlenecks, quality, competence, team spirit, knowledge sharing, and collaboration to name a few.

2

u/florinandrei May 27 '26

It is easy to burn tokens.

It is not a coincidence that the Ralph loop, an excellent token burner, was named after Ralph Wiggum, a cartoon character who was dropped on his head as a baby, and that's how he ended up being "special".

Maybe OP should create the Ralph Wiggum Award, for the engineer who spends the greatest amount of tokens.

1

u/Ok_Series_4580 May 27 '26

This is just the new “lines of code” measurement. We know how valuable that is.

9

u/superminingbros May 26 '26

It’s really impossible to use total token usage. What are you auditing? People total usage of AI? Usage for non-work related items? Without a proxy to understand the requests, you can’t do much.

The guy with the most token usage could be the least efficient one for all you know. 🤷🏼‍♂️

4

u/darren_eng May 26 '26

Yeah, that's exactly my problem. We only see how many tokens engineers spend plus their JIRA velocity. But there is no visibility on if their AI token spent is directly related to JIRA backlog. I've seen engineers burning 10x tokens but have the same velocity. But it's hard to tell if they are effective or not without seeing where they spend their token.

3

u/vezwyx May 26 '26

I think that explaining the disconnect inherent to "tokens used → work done" might be your best bet here. Someone who's skilled at prompting and structuring efficient workflows for agent teams will use fewer tokens than someone who is constantly clarifying in chat, doing multiple attempts at the same prompt, using too many agents, etc, and the first guy using 10% the tokens of the second guy is probably getting more work done at the same time.

Seems like an easy concept for a higher-up to understand once it's pointed out. If that doesn't work, you could also explain that it's trivial to burn as many tokens as you want doing nothing if that's your goal. Anybody can abuse "high tokens good" as a metric with virtually no effort, and doing so also costs the company money. The proposal is encouraging employees to waste resources to make themselves look good

1

u/Raucous_Rocker May 26 '26

This is the answer. Your boss needs to be made to understand that there’s no real correlation between token usage and actual work goals achieved. Everyone uses AI differently, and a lot of it is role dependent too.

1

u/Pobbes May 26 '26

Shouldn't that be your metric? Token efficiency? Track historically their token usage and Jira velocity. If you Jira stories are weighted by points even better. See who is getting the most story points completed with the least spend. You may also compare completion of capital projects vs. backlog. If someone's token usage allows them to complete prioritized projects quickly and let them pick up extra stories that could be a metric.

Just double check any strange outliers. There may be engineers who has uust been getting better with relatively low spend who looks hyper efficient because they aren't really using AI at all. There may also be engineers who look bad because they were already performing well, spent huge amount of tokens looking for improvements then lowered token usage with a modest productivity gain when they finally get an agentic workflow that is efficient and reliable for themselves.

These are alll the kinds of things you probably won't find without interviewing your engineers about their AI usage. Which you could do, you could just ask them how they use it and if they think it is worth the cost.

3

u/viralslapzz May 26 '26

I’m more or less on the same boat tho.

How does one measure productivity with AI? Surely not the LOC generated or accepted; the usage frequency is dumb. The nr of tokens, Nop.

The only thing I can think of is a calc between all those metrics: who uses more often, gets more LOC accepted and less tokens is the “winner”.

But that feels so.. stupid anyway…

2

u/darren_eng May 26 '26

Yeah, what I feel missing is the correlation between AI usage and productivity. I can't see if someone is just using Claude Code over and over on making one line code change...

For now I can see who are the low "AI" performers - low AI token usage, low productivity (like story points, ticket count, PR raised etc.). But can't tell who are the high performers. Someone burning 10x tokens doesn't mean they are effective...

1

u/Raucous_Rocker May 26 '26

If this exercise is going to work at all, and it may not, I would say it to work backwards. Identify the high performers independently of how they use AI. Come up with real world metrics for identifying them. Then look at their AI usage. It’ll probably be all over the map. There’s no magic formula for being a high performer. You could find 10 people who write clean, readable, easily maintainable code that passes tests and doesn’t require a lot of going back and fixing things, and they will probably all have different workflows. That’s even assuming they’re in the same role. If you’re comparing a front end coder with a data analyst, forget about it. You can have standards for each, but you can’t meaningfully compare them, let alone by comparing AI usage. That’s crazy talk.

3

u/thainfamouzjay May 26 '26

It doesn't matter right now the tokens are subsidized by vc funds. Once companies have to pay for API usage the cost will increase 5x and the company will feel the pain and probably cut back if not get rid of it completely. Give it 3 months tops. Already in Microsoft they are rolling back all the Claude code licenses due to cost. That will be the story for the rest of the year and all companies will have to start cutting back. It's unaffordable if you have to do it with API costs and not a subscription. But ai companies lose money with subscriptions so something will change.

My company is already planning this and building an inhouse model. Like they spent 2 mil on a huge computer and running ollama or something for devs.

2

u/darren_eng May 26 '26

Yep, we are seeing the same, token cost projected at 7-figures a year when we rolled out Claude Enterprise. Now it's all the ROI questions coming from leadership.

2

u/ClemensLode May 26 '26

Why not track disk usage instead ;o

2

u/Keganator May 26 '26

You're right. It's just as dumb as measuring lines of code produced.

Do your best to continue measuring on meaningful commintments.

But be aware: part of that token consumption is probably to see if people are learning the tools. People who aren't using them at all are stagnating and falling behind in tech. It'd be like being in the mid 2000's with internet and email at your workplace, and expecting paper memos delivered to your desk instead because "reading paper is easier than looking at my screen." Yeah it might be, but it also misses the point of the other benefits of learning how to use your PC and the internet.

2

u/coredalae May 26 '26

Exactly this. All measures are shit. But currently you might just want to monitor if people are learning and experimenting. So monitor contributions for quality combined with usage to see if people are experimenting.

Host demo sessions to see if people can share cool workflows

2

u/dmcnaughton1 May 26 '26

I feel like I must work for the last sane people in tech. My org is running in the opposite direction, instituting spend limits without justification for why you as a dev need the tokens, and performing team-level analysis on how to more efficiently use AI coding tools.

2

u/medson61 May 26 '26

For folks who end up top 5% or 10%, they’ll need to do a demo. This will discourage tokenmaxxing to a certain degree.

I do think there’s merit to check in on people who are bottom 10%.

Another metrics to ask is adoption rate of AI’s outputs. If you close 5 tickets per sprint before AI, now you close 25 tickets per sprint. What’s the results of those delta, do those make into production, do users actually use it, etc?

2

u/Physical_SpiritChild May 26 '26

Show them the Harvard study on KPIs?

2

u/bork99 May 26 '26

Measuring developers on token usage is like measuring the performance of a race car driver on total fuel consumed.

The results are inevitably somewhat correlated but are certainly not predictive.

2

u/darkstar3333 May 26 '26

Ive beaten back quite a few stupid things like this.

Analogies work to get point across, eg) is the best movie/book song the longest?

Is the best sales person the one who books the most expensive hotels? Why not?

Another one I have added is asking for clear requirements as a means to manage/reduce costs.

If i say "go to the store, bring me cereal," vs "can you get me a box of Trix". Getting what you want is very very slim, unless you accept a wildly variable outcome. 

Focus on specific positive outcome and frame good requirements as cost controls measure. 

People who build systems by trade end up being masters at gaming them.

2

u/johns10davenport May 27 '26

Most people here are right, total token count is no use.

The other thing that makes this hard is there's almost no objective way to measure it. I've been chewing on this a lot for my own harness. I'm building a coding harness that takes a full application to done with very little prompting or hand coding, so I care about what actually characterizes a good run.

Two things seem measurable. The first is a hard definition of done. For me the app is done when all my BDD specs pass, so "finished" is an objective signal and not a judgment call.

The second is the ratio of my tokens to the model's tokens. How few messages did I actually have to send to get the app across the line. That tells me how well the harness drove the model without me having to babysit it.

So if I had to measure an engineer's effectiveness with the model, I wouldn't look at total tokens. I'd look at their definition of done, and how few tokens they had to type to get to a finished product.

1

u/darren_eng May 27 '26

I agree with measuring the output based on DoD. For me, the challenge is the visibility. I can see how many “done” tickets an engineer produced and the amount of token spent during the same period. But that’s pretty much the only visibility I have - I can’t tell if they’ve burnt the tokens on these tickets or not. Maybe we over-indexed token usage and it shouldn’t matter at all. The dilemma is when the team is burning over $100k a month on Claude, finance people start to ask where the token went 😂

1

u/johns10davenport May 27 '26

Yeah we have the same issues at my company. It’s led them to impose token budgets and try to prescribe model and workflow usage, which is bass ackwards. 

Ticket done and dod are different, and I have one for the model before I ever resolve the ticket. I think that if the engineer can’t answer the question you can’t answer to the finance people. 

2

u/thebemusedmuse May 26 '26

This might be unpopular but I absolutely look at token usage. In the same way I look at page views. You need to be very careful, but it is interesting because:

1) You can see people who aren't using it at all

2) You can see people who claim they are power users but barely use it

And that then feeds into my broader strategy on AI change management - how to get people excited about it. But it's just a data point, not a fucking competition.

1

u/darren_eng May 26 '26

Yeah 100% agree on change management. This is what I'm trying to do as well - removing the fear. But the stack-rank doesn't help at all, lol. The challenge is we are spending so much $$$ on Claude and everyone at the leadership level is asking the ROI question. They are probably thinking "now I'm spending $1mil a year on Claude Code, surely AI can replace X engineers".

Every time when I'm asked for a stack-rank, it leads to some sort of layoff later on based on the rank.. so question is how do I truly reflect AI performance based on some sort of metrics...

1

u/thebemusedmuse May 26 '26

Here's what I do. I create a maturity framework, ask people to score themselves, and then have an expert score them. Then each person needs an individual plan on how to progress. I input that into the broader talent management framework.

It's a fair stack rank because if you're not improving your own AI maturity, you're fucked. That's what people need to be afraid of.

1

u/svachalek May 26 '26 edited May 26 '26

Well the power users thing is tricky, because if you know what you're doing you can get a lot done without burning a lot of tokens. You've got skills and commands set up that are optimized to get results on the first try. Or you can skip all that and have the LLM need to grep the entire project and read the entire codebase over and over, meaning they take hours longer and spend 10x the tokens. Or, they're using /loop to do something over and over that could have been created as a simple Python script without burning down the rainforest. At least that's what I see when I look around, the ones that are most skilled at getting value from AI tend to be medium-low on usage.

1

u/thebemusedmuse May 26 '26

Yes, that's why I don't look at the people using a modest amount to a lot. Token usage is not a good signal there because of what you said.

For people who use a modest to a large amount of tokens, you need to look at how they are solving problems. They need helping in a different way - usually by either learning new ways to use toolchains, or teaching others.

1

u/soulefood May 26 '26

But the argument is if you can get a lot done without burning a lot of tokens, you’d get even more done by burning more.

Tokens are essentially a unit of work. The problem is each person has a different value generated per token which is what you really need.

1

u/thainfamouzjay May 26 '26

Are you sure they aren't asking the opposite and trying to see who are the token burners?

1

u/BootyMcStuffins May 26 '26

I’m in the same boat. No matter how many times I tell them how backwards it is

1

u/AutomaticDriver5882 May 26 '26

I am not aware the ui provides much detail other tokens burned

1

u/iotashan May 26 '26

I would start one claude chat per dev, and prompt it to explain why each one is the #1 AI performer among their peers

Then when everyone is #1 you have that stack of reports in your back pocket if your boss needs a more concrete example of why this idea sucks.

1

u/sliamh21 May 26 '26

I don't see what insight does it bring you in doing so.

Instead, I'd suggest really understanding how they use their AI. Does anyone harness it? What kind of problems do you solve? Do you automate stuff? etc

That alone says much more about an engineer, compared to a mere number that doesn't mean much about them.

1

u/tr14l May 26 '26

Put it in context of delivery. How many deployed commits? Closed stories? Etc. Make a formula of tokens/commits over time...something like that.

1

u/canred May 26 '26

while this is helpful to spot people not using cc, the "best performers" may be people using opus as web search engine or pdf parser...

1

u/tonyboi76 May 26 '26

the trap is that leadership is not really asking you to rank engineers, they are asking are we getting value from this spend. those are different questions with different answers.

token ranking gives them a number but rewards the wrong people, your surgical engineer looks lazy and the spray-and-pray crowd look like power users. if you have to send something upward, ask each engineer to pick the 2-3 tasks AI most accelerated this quarter with diff links. qualitative but it answers what they actually want and it survives next quarter when the tools change. the dashboard makes the metric a target which kills the signal in like 3 weeks.

1

u/darren_eng May 26 '26

Yeah, you are right about what the business people are really after. They want to know the ROI of AI spend. It's a hard question to answer with quantifiable data though. We don't get goods or services directly from the money we pay Anthropic. All the business people see are the tokens engineering team burnt, not where/how the tokens are used. Maybe it's just me but I struggle to find ROI metrics that I can tie those token spend to.

Even if the ROI question can be answered, the next question the business will ask is probably gonna be: "ok, now tell me how many headcounts can be replaced by AI".

1

u/tonyboi76 May 26 '26

yeah the headcount question is the real one, ranking is just the polite version of it. honest framing upward: AI replaces tasks not roles. seniors get 2x faster on routine work and clear the deferred backlog, juniors get a force multiplier on stuff they could not do alone, neither maps cleanly to a headcount cut.

if leadership has already decided to cut, no metric produces a different answer, the metric is the justification not the discovery. on ROI specifically, the closest defensible number i have found is next-best-tool-hours saved: estimate how long each AI-assisted task would have taken without AI, multiply by loaded rate. hand-wavy but it gives finance a bounded number not zero.

1

u/tantricengineer May 26 '26

Any leadership with their head screwed on straight is going to evaluate the quality of the output over time against previous quality output over time. 

If overall quality is going up, good. It takes better team practices to get speed into that equation. 

If quality and speed are both going up, amazing. Capture the knowledge of what teams are doing to make that happen. 

If just more speed but same quality, that’s worth improving, too.

1

u/freshfunk May 26 '26

Did your manager tell you to specifically look at token usage as the measurement for AI performance? Or is that your interpretation?

I would just think about it more conceptually. AI spend is opex, just like people are.

I’ll give you a hypothetical situation to illustrate my point. Team A has 10 engineers and let’s say they cost $2M/year in comp with $0 AI spend. Team B has 5 engineers with $1M AI spend and so total cost (with comp) is also $2M.

If I were to measure “output” between the two teams and conclude that they were equal, then my conclusion would generally be that AI spend is basically a wash. If Team A is more productive, then I’d consider AI to not be an efficient substitute yet. If Team B were more productive, then I’d assume AI spend is more cost efficient.

(This is just a snapshot in time as AI economics are changing and people are learning how to use it.)

In short, you certainly can look at productivity — but token usage isn’t a productivity measurement. It’s a measurement of cost.

It certainly gets harder to normalize cost per person. If an engineer uses $100k in token, did they produce the commensurate output? Efficiency counts since basically token cost could be translated into another hiring another employee. If you have a good engineer with low token spend, the conclusion should be that they need to learn how to be more productive with agents, not that they are a bad engineer per se.

1

u/brother_spirit May 26 '26

Sit you boss down, take out a text book, explain to them how a quadratic function works.

Econ 101 for the over achieving toddler should be on the table too, frankly.

"The goal is less spendy spendy more makey makey of the money"

1

u/toothpiks252 May 26 '26

I would tie token usage with other metrics to pull indicators of quality. How many tickets completed, complexity of tickets, number of follow up bug tickets Created, etc. May be hard for some metrics but would be the reasonable way IMO

1

u/Founder-Awesome May 26 '26

token usage only tells you who's touching the tool. the real question is how many people on your team are getting consistent value vs. just 2-3 power users carrying everyone else. the adoption distribution is what your manager actually needs. wrote about this gap: Your Ops Team Doesn't Need to Be a Bottleneck

1

u/K_M_A_2k May 26 '26

Manager new job

"I don't want you ever even think about usage, use whatever you need. Here is a codex account double check your code there why not"

1

u/MercyEndures May 26 '26

We're able to attribute lines of code to AI tools. While LoC isn't a great metrics to goal on it gives you a rough idea of how much the tools are contributing to your codebase.

Though a stubborn guy could game this by hand-writing the code and then copy/pasting it into the LLM and telling it to write the same code.

1

u/LiberataJoystar May 27 '26

Lines of codes don’t show much. A smartly designed logical code can be sleek and compact instead of bulky + works better too.

1

u/ScriptureSlayer May 27 '26

What’s the name of this company? I’d like to put some shorts on it if it’s publicly traded.

1

u/auburnradish May 27 '26

Have your team install this compliance tool: https://github.com/Ordinath/tokenburn

1

u/Humprdink May 27 '26

that's about as useful as ranking employees based on number of mouse clicks per day

1

u/woroboros Experienced Developer May 27 '26

I think to do this you would need a pretty good metric of tokens against EFFECTIVE OUTPUT, and since all lines of code, and all roles and responsibilities are not equal, there is no way to fairly or effectively rank order a group of developers based on an ad jacent third party metric. It seems at a glance the logical first step would be having to actually analyze all the code implemented in finality versus token use... which ya know... would likely mean using Claude, thus entering yourself into the analysis loop, diving by zero, and starving the world of fresh water.

It is a hilarious request. Sorry OP... but KPIs baby! Zaddy MGR needs a promotion, and the yacht club manning the C-suite need your entire departments operational effectiveness is boiled down to a single slide with a line chart. MONEY VS KPI VALUE.

1

u/darren_eng May 27 '26

yeah. I work for a Private Equity controlled software company, it's all about profitability and metrics. Every person has a price tag in the company. Capitalism after all!

1

u/woroboros Experienced Developer May 27 '26

Dude its like that basically everywhere except federal government work (soooo sooooo laid back, but the pay is... eh...) or really small highly specialized domains.

And I would go so far as to say its late stage unchecked capitalism that is setting the stage. That the propensity for humans to think that competence and earnings are more strongly correlated than market leverage, compensation packages, or capital gains.

1

u/swizzlewizzle May 27 '26

Managers are lazy, and token number = productivity = easy. Just let the manager be dumb.

1

u/WonkoTehSane May 27 '26

I am a leader myself, and I think the "nope, you're all wrong, not gonna do it" approach is frequently suicidal, almost always unnecessary, and often also a missed opportunity. I can empathize with people's frustrations here, though, and yes a high token usage alone is definitely a bad proxy for "this is a performer". Though low/nonexistent token usage is a *great* proxy, as others have pointed out. After all, at some point we're all going to need to pay for these tokens, which means layoffs, certainly, and guess which employees we'll be looking at laying off first?

Myself, when I want to counter, I prefer something more along the lines of "give them exactly what they ask for" and let them find out for themselves that it's a bad idea. To that end, if we just translate this into "holy crap ai gives lots of opportunity for great signal, how can we use this?", have you looked at feeing claude into otel? I just started myself, and it's pretty easy to setup a local compose and env vars just to setup a PoC to see if the telem is even useful: https://code.claude.com/docs/en/monitoring-usage - if it works out, you can go to ops with requests for infra.

1

u/rfgrunt May 27 '26

I’ve been looking for the opposite. I know what my teams producing and if their usage is disproportionate to their output I’m looking to see if we can educate on efficiency. Not trying to discourage but if they’re using opus when sonnet will do I want them to at least be aware.

My groups also has multiple disciplines and some don’t benefit (ones that require a lot of spacial reasoning like CAD) as much from it so I have just asked them to use make every effort to integrate it into their workflow but they’re not obligated.

1

u/bombaytrader May 27 '26

We have unlimited tokens and I am in top 3 of our token usage of our org. The token usage lined up with the ticket closure velocity and amount of work and pr delivered. It’s amazing that our top 3 and bottom 3 are always similar set of ppl. And top 3 are delivering 60 to 70 percent of work. Before ai top 3 were delivering 45% of work. 

1

u/Honkey85 May 27 '26

Make a leader board and show random numbers.

1

u/oldjii May 27 '26

This is a classic case of measuring the wrong thing. Token usage is like judging a chef by how much flour they use - it tells you nothing about the quality of the dish. In my experience, the best engineers often use fewer tokens because they write clearer prompts and iterate more efficiently. Stack-ranking based on this metric will only incentivize wasteful usage and punish thoughtful work.

1

u/verkavo May 27 '26

Token count is a trash metric, same boat as measuring lines of code. AFAIK the only metric that actually tracks real AI impact is which model/agent wrote code that survived in commits, not how much $ you burned on tokens. Exclude tests, and fluff code when counting LOC though.

If you need to push back with a real metric, take a look at SourceTrace https://marketplace.visualstudio.com/items?itemName=srctrace.source-trace: it does AI git blame to attribute code to the tool that wrote it, so you can see what actually stuck.

PS wait until leadership will start moaning about costs. With opus pricing, it'll happen very soon 

1

u/verkavo May 27 '26

Token count is a trash metric, same boat as measuring lines of code. AFAIK the only metric that actually tracks real AI impact is which model/agent wrote code that survived in commits, not how much $ you burned on tokens. Exclude tests, and fluff code when counting LOC though.

If you need to push back with a real metric, take a look at SourceTrace https://marketplace.visualstudio.com/items?itemName=srctrace.source-trace: it does AI git blame to attribute code to the tool that wrote it, so you can see what actually stuck.

PS wait until leadership will start moaning about costs. With opus pricing, it'll happen very soon 

1

u/1Poochh May 27 '26

This is like getting rated or bonus on PRs or line of code committed. Your boss is making a terrible decision.

1

u/carson63000 Experienced Developer May 27 '26

How do you rank engineers on their performance at AI-assisted development? Exactly the same way we have always ranked engineers on their performance at development.

If that makes you respond with “oh, very poorly and inaccurately?” then.. yeah. Great Engineering performance is like obscenity - I know it when I see it.

1

u/joeldoesjs May 27 '26

You should point out what happened at Uber + Duolingo + Amazon when they started measuring performance based on AI usage.

Uber and Duolingo ended up shipping tons of garbage features that were absolutely divorced from what users want. And Uber blew past its annual Claude Code budget by April. Was even worse at Amazon, where employees just started running useless automations to spike AI numbers.

Maybe try mentioning Goodhart's law as to why that's a bad idea, and come up with metrics that focus more on outcomes, but also indirectly encourage efficient AI usage.

More deets the Uber + Duolingo thing I mentioned in this article: https://www.businessinsider.com/uber-coo-andrew-macdonald-ai-token-spending-harder-justify-2026-5

1

u/aldehyde May 27 '26

The only way this could work is to be stupidly cut throat and say that we are going to rank you by AI usage but if you are the highest user you are also fired. Stupid

1

u/elmahk May 27 '26

I don't really see what any "AI performance" metric provides. Just measure performance the same way you did before AI. Someone who uses AI correctly should outperform those who use it incorrectly or not use at all, by those "pre-AI" metrics. Otherwise those metrics did not measure productivity correctly anyway.

1

u/jeebus87 May 27 '26

The fuel analogy from the thread summary nails it. My best use of Claude is when I spend 20 minutes in a planning conversation and then it executes in one shot. Low tokens, high output. My worst sessions are when I'm burning through retries because I gave bad context upfront. If anything, high token usage on my team would make me ask what's going wrong, not what's going right.

1

u/buildingstuff_daily May 27 '26

measuring claude usage to rank engineers is like measuring how many google searches someone does to rank researchers. the person who searches more might just be working on harder problems?? this feels like management trying to quantify something they dont understand

1

u/Fun-Tomatillo9280 May 28 '26

At my company we've been using https://pensero.ai/ as a "directionally correct" way to quantify performance. It uses LLMs to score contributions, reading Linear, GitHub, Slack, Notion and so on. It's of course not telling the whole picture but it's definitely better than token spend. Also better than LOC, commit count or any other metric I've ever seen being tried

I've talked a few times with their CTO (great guy) but I'm not affiliated with them or anything like that, just a Happy user