Messages sent, active users and self-reported hours saved can show adoption. They do not prove productivity. A team may draft faster but spend longer checking errors, create more work than reviewers can absorb, or use saved time for low-value output.

To know whether ChatGPT or Claude helps, measure a defined workflow from start to approved business outcome. Include quality, review, failure and what the team does with the capacity released.

Choose a unit of work

Avoid measuring “marketing productivity” or “engineering productivity” as one number. Select repeatable tasks such as:

  • Approved campaign brief
  • Sourced market-research memo
  • Bilingual article ready for publication
  • Customer-support response approved for sending
  • Tested software change accepted in review
  • Weekly performance report with verified metrics

Define the starting point and finish. Time ends when the output is approved, not when the AI produces a draft. Record volume, complexity, quality threshold and responsible reviewer.

Baseline the current process

Before the AI pilot, sample enough normal tasks to capture variation. Measure:

  • Hands-on work time
  • Waiting and handoff time
  • Review and correction time
  • Rework and rejection
  • Error or defect rate
  • Throughput and backlog
  • Business outcome where observable

Use actual timestamps or structured time sampling where practical. People estimate time poorly, particularly for fragmented work. If a baseline is unavailable, run AI-assisted and normal processes in parallel for a limited period.

Do not compare the strongest AI enthusiast with an average non-user. Account for experience, task difficulty and starting quality.

Measure quality-adjusted completion time

The central metric is:

quality-adjusted time = creation + prompting + tool runtime + verification + correction + approval + expected defect recovery

Count a task complete only if it meets the same acceptance standard as the baseline. For writing, check facts, sources, brand and bilingual meaning. For code, run tests and review correctness and security. For analysis, verify calculations and decision usefulness.

Track first-pass acceptance and final acceptance separately. Faster first drafts with lower first-pass acceptance reveal work transferred to reviewers.

Separate speed, output and value

AI can affect work in at least four ways:

  1. Time compression: the same approved output takes less time.
  2. Quality improvement: the same time produces a better approved output.
  3. Throughput growth: more approved work is completed.
  4. Task expansion: work that was previously uneconomic now gets done.

Anthropic’s internal research reports both self-reported productivity and increased output, while noting measurement limits and active supervision. It also reports that some AI-assisted work would not otherwise have been completed. That matters: time saved alone misses new work, but new work is valuable only if it supports a real objective.

Track each effect separately. Do not turn extra low-priority output into a productivity claim.

Treat usage analytics as diagnostic

Workspace analytics can show activation, active users, messages, tools, projects, skills and task patterns. They help identify adoption gaps and candidate workflows.

OpenAI’s documentation explicitly notes that workspace impact signals are aggregated and self-reported and do not represent causal ROI measurement. High usage may indicate value, confusion, repeated failures or enthusiasm. Pair it with task results.

Likewise, a low-use expert may apply AI to one high-value monthly decision. Messages per user would undervalue the contribution.

Add outcome and guardrail metrics

Connect the task to an outcome:

  • Marketing: qualified conversion, campaign learning velocity or content-influenced pipeline
  • Sales: response time, accepted opportunities, meeting attendance and revenue
  • Product: time from evidence to tested decision, adoption or retention
  • Engineering: accepted change cycle time, escaped defects, reliability and incidents
  • Operations: turnaround, accuracy, backlog and cost per case

Add guardrails: privacy incidents, unsupported claims, Arabic corrections, customer complaints, security findings, reviewer overload and employee satisfaction. A speed gain that raises material risk is not a net gain.

Run a controlled pilot

Select a representative group and task set. Where practical, randomise tasks or alternate periods between AI-assisted and baseline processes. Keep acceptance rules, reviewers and input quality consistent.

Capture at least:

  • Task ID, type and difficulty
  • Person and experience band
  • Tool, plan, model or mode
  • Prompt or workflow version
  • Start, draft and approval timestamps
  • Human review minutes
  • Corrections and severe errors
  • Approved outcome and business result
  • User-reported effort and satisfaction

Blinded quality review reduces preference bias. For bilingual work, use qualified English and GCC Arabic reviewers.

Beware of perceived productivity

People can feel faster while taking longer. METR’s randomized 2025 study of experienced open-source developers on familiar repositories found a gap between perceived and measured effects in that setting. The result should not be generalised to every worker or newer tool, but it demonstrates why self-report needs observed task data.

Conversely, laboratory tasks may miss benefits from better coordination, more experimentation or work that becomes possible. Use multiple methods: operational telemetry, quality review, surveys and business outcomes.

Calculate net capacity and ROI

For each workflow calculate:

net hours released = baseline quality-adjusted hours − AI-assisted quality-adjusted hours − enablement and admin allocation

Then ask where those hours went. Capacity used for customer work, additional experiments, backlog reduction or improved quality may create value. Capacity absorbed by more meetings or unchecked content may not.

Estimate:

net value = value of released capacity + incremental outcome value − licences − integration − training − review − incident and error cost

Use ranges instead of false precision. Separate observed results from assumptions, and state the evaluation period.

Report distribution, not only averages

Show median and spread by task, team, experience, language and complexity. Averages can hide that experts benefit on routine work while novices need more review, or that English improves while Arabic corrections rise.

Review severe failures separately. One data leak or harmful customer message can outweigh many small time savings.

Decide what to scale

Classify workflows:

  • Scale: clear quality-adjusted time or outcome improvement with controlled risk
  • Improve: promising but review, context or training remains costly
  • Limit: useful only for specific users or cases
  • Stop: no net value or unacceptable failure

Document the configuration, training, review and controls behind success. Buying more licences does not reproduce a workflow automatically.

Re-measure after model, process or team changes. Productivity is a property of the combined system: people, model, context, tools, controls and incentives.

Use a balanced executive dashboard

Report five lines per workflow: adoption, approved throughput, quality-adjusted time, business outcome and risk or rework. Add a short explanation of what changed and what decision is requested.

Do not promise an enterprise-wide percentage from a small pilot. Scale the evidence with the decision.

ChatGPT and Claude may save substantial time, improve work or enable new tasks—but only workflow-level evidence can show where. DEMA can build the baseline, pilot design and bilingual outcome dashboard for your team. Request a free growth audit or book a free consultation to turn AI adoption into measurable operating value.

Sources