Performance management has been broken in most organizations for a long time.
It was designed for offices, where managers could track progress through proximity and ambient visibility. Remote work removed the ambient visibility. AI is now removing the need for most traditional performance proxies. What is left, the actual work of understanding whether someone is doing good work and having a meaningful conversation about it, turns out to require more clarity and less administrative overhead than most companies have built.
At Running Remote 2026, Nadia Vatalidis, Head of People at Doist, and Nick Francis, Chairman of Help Scout, shared what that looks like in practice. Moderated by Shelby Wolpa.
The 90/10 principle
Nick Francis has spent time watching his own thinking evolve on this. His observation: AI can now handle roughly 90% of the synthesis work in performance evaluation.
Aggregating peer feedback across six to twelve months. Summarising goal progress and outcomes. Identifying patterns that a human manager’s brain struggles to hold simultaneously, and addressing the recency bias that has always distorted performance reviews, where the last three weeks dominate the last year in most managers’ minds. AI remembers everything. Human memory does not.
The last 10% is what requires humans. The final accountability. The quality of the actual conversation. The creativity and craft in interpreting what the data means for a specific person in a specific context. The relationship knowledge that no system can hold.
Hard conversations and expressions of genuine recognition must happen between people who have thought carefully about what they want to say. Using AI for these interactions, both leaders agreed, would be soulless and counterproductive. Employees detect it immediately, and it signals that their manager does not actually care.
Doist’s 48-hour review cycle
The practical demonstration of the 90/10 principle was Doist’s February performance review cycle.
105 people. Self-reviews and peer reviews. AI summarization of all peer input into final summaries. Complete cycle: 48 hours. The typical comparable process at most companies takes six weeks.
Each employee completed their own review and a maximum of five peer reviews. AI aggregated and summarised peer input, allowing managers to focus on the qualitative last 10%, the actual conversation with each person about what the data means for them.
What made this possible was not the technology. It was the performance philosophy established before any framework was built. Doist’s career framework, publicly available in their handbook, separates levels from job descriptions. Peer reviews are not anonymous, Doist operates under a radical candor norm where direct feedback is both safe and expected. The framework has been through seven or eight major iterations based on employee feedback.
The sprint philosophy behind the 48-hour window: more time given equals more procrastination. Compress the window, pair accountability with the constraint, and people produce better work with less anxiety. The review is an event, not a season.
Outcomes, not effort
Both companies have made a deliberate choice to define high performance in terms of outcomes and business impact, not effort or activity.
At Help Scout, effort is explicitly not a performance dimension. If time and energy are not producing the right outcomes for the business or customers, that is a signal to pivot, not a signal to appreciate the effort. This is a different psychological contract than most organizations operate under, and it requires significant clarity in how expectations are communicated.
At Doist, a 105-person company, every role has measurable impact on the business. The bar keeps rising, not because of AI, but because the definition of what is achievable keeps expanding.
For HR leaders, this is a meaningful design choice. An effort-based culture can produce activity without impact. An outcomes-based culture requires more clarity upfront, about what good looks like, how it is measured, and what decisions get made when outcomes are not being reached, but it creates the conditions for genuinely high performance.
What to measure: ARR per employee
Help Scout’s performance conversation extends beyond individual reviews to how the organizations measures its own performance.
The primary metric at board and C-suite level is ARR per employee: annual recurring revenue divided by headcount. The goal is to continuously increase this without burning people out or reducing headcount through layoffs.
Doist maintains extremely low attrition alongside this ambition, more than 50% of employees have been there five or more years, and 21 people celebrated ten-year anniversaries in 2026.
On AI metrics specifically: performative numbers, lines of code generated, percentage of code written by AI, AI session counts, are dismissed by both leaders as meaningless distractions. The question is whether AI initiatives are moving core business KPIs. Everything else is noise.
What does AI actually save your team? Now you can measure it.
Time Doctor’s AI Impact Report gives managers visibility into AI adoption across groups, estimated time saved, and business value generated by role. You can see which tools your team uses most, how usage breaks down by task type (writing, research, code generation, and more), and whether prompt quality is improving over time.*
Explore the AI Impact Report →
*Available to Advanced Data Capture customers
Where human accountability is non-negotiable
The most memorable moment of the session was a cautionary example.
A manager copy-pasted a ChatGPT response to an employee’s sick message, complete with the prompt fragment “if you’d like me to add.” The employee saw it. The trust damage from that single moment was disproportionate to the effort saved.
Using AI to respond to a sick message is not efficiency. It is evidence that the manager does not actually care about the person who is sick. And employees know the difference.
Doist uses Gemini Gems AI coaches to improve feedback quality before writing reviews, using AI as a soundboard for preparing meaningful conversations. That is genuinely useful. AI as a replacement for the conversation itself is not.
The principle holds across all meaningful manager-employee interactions: AI belongs in the preparation, not the moment. The quality of the moment is the whole point.
What this means for HR and people leaders
The picture that emerges from Doist and Help Scout is of performance management rebuilt from first principles: clear on what matters (outcomes), honest about what AI can and cannot do (synthesis yes, relationship no), and deliberate about the culture required to make it work (radical candor, psychological safety, genuine accountability).
For distributed teams, where the ambient visibility of an office has never existed, this is not a future state. It is a requirement. The managers leading distributed teams well right now are the ones who have real clarity about what their people are working on, how that work connects to business outcomes, and what a meaningful conversation about performance actually looks like.
Time Doctor gives people leaders the underlying visibility into how work actually flows across their distributed teams, the data layer that makes an outcomes-based performance conversation possible. Not to watch people. To understand the patterns that inform better decisions about how to develop, support, and lead them.

This blog post draws on the session “Humans + Machines: Unlocking Performance in Remote Teams” at Running Remote 2026 in Austin, Texas. Featuring Nadia Vatalidis (Head of People, Doist) and Nick Francis (Chairman, Help Scout), moderated by Shelby Wolpa (Shelby Wolpa Consulting).

