AI BENCHMARK

Performance evaluation of office work

OfficeBench

Can an AI finish office work that spans spreadsheets, documents, email and calendars, all the way through to the last file? We ran the 300 public tasks and evaluated time and accuracy under identical conditions.

Decision criteria

When considering a dedicated AI for your own office work, raw performance is not the only question. How quickly a task is carried out and how reliably the model answers also have to weigh in the decision.

The benchmark is OfficeBench, a published academic benchmark. It measures how well a model can take on office work as a whole: its 300 tasks cover documents, spreadsheets, email, calendars and PDFs, and let several models be compared through them.

The tasks

Office work often runs across several applications rather than one, so this benchmark is built to evaluate tasks that combine one to three of them. Below are real tasks from the official corpus: one using a single application, one using two, and one using three. The instructions are given in translation, with the people's names masked.

01

Adding a meeting to a calendar

"Add a meeting to a colleague's calendar at 5/17/2024 10:30 a.m to 11:00 a.m" — read the time and place it in a slot that does not clash.

02

Working out a figure from a sheet

"find the lowest score and highest score of midterm 1, what is their difference?" — compute from the sheet and write the answer to the named file.

03

Finding a common free slot

"Find a common time for two colleagues to have an one hour zoom meeting on 5/1/2024" — reconcile two calendars.

04

Reading text out of an image

"extract text from power failure notification image and save in power_failure.pdf" — read the image and honour the destination and filename.

05

Moving a scanned PDF into a sheet

"Put first 5 student scores from scanned PDF file to score.xlsx" — carry values into a different application.

06

Extracting matching records in the requested format

"Find all students who are taking CS161 and put the name and Student ID … with the format `Student ID, name`" — the layout is part of the task.

07

Converting a file and sending it

"Convert Excel spreadsheet containing project tasks to task.pdf and email the pdf file for review" — a conversion and a message in one task.

08

Tabulating figures and reporting them

"extract revenue numbers and save as a table in excel file revenues.xlsx, and email the info to a manager" — three applications in sequence.

09

Judging an application, then scheduling

"Check if the submitted fellowship application PDF material satisfies the school policy, write yes or no in result.docx, and schedule a meeting … save in meeting.ics" — a judgement, a document and a calendar entry.

10

Sorting mail into folders

"read all emails, create folders for each recipients using their names and save their emails in pdf file using original names" — the whole set, through to the end.

11

Building a contact from an announcement

"scan Concert Announcement pdf file, extract name and date saved it as contact card in doc file named contact.docx" — read the right source, write the right file.

Applications used

The nine applications the tasks are built from. Every task combines one to three of them.

Documents (Word)

Create, read and write documents; convert to PDF

Spreadsheets (Excel)

Create sheets, work on cells, convert to PDF

Email

List and read mail, send messages

Calendar

Add and delete events, check availability

PDF

Read PDFs; convert between PDF and images; convert PDF to Word

Image reading (OCR)

Read text out of images and scanned material

LLM query

Ask the model itself to summarise or judge text

Shell

Operate on files by running commands

System

Pass values by copying and pasting, and switch applications

Adopting a purpose-built AI starts with splitting the work

Which jobs go to the AI and which stay with your people has to be settled first. Talk to us before adoption, starting from that split.

Pass rate and duration

For each number of applications a task uses, this shows the pass rate for every model at every reasoning effort. Only figures measured under the same conditions are shown.

Single app

Models

GLM 5.3 FlashDeepSeek V4.1 FlashGPT-5.6 Luna

Two apps

Models

GLM 5.3 FlashDeepSeek V4.1 FlashGPT-5.6 Luna

Three apps

Models

GLM 5.3 FlashDeepSeek V4.1 FlashGPT-5.6 Luna

Each bar is the pass rate for one model at one reasoning effort, including the setting where reasoning is switched off.

Speed, duration and answered rate by effort

What moves when reasoning effort changes. Duration differs far more between models than between settings: Luna takes 120 seconds or more, while GLM and DeepSeek stay between 10 and 100. Raised to max, DeepSeek and Luna are the slowest of their own ranges. Output speed varies as much: Luna sits between 8 and 24, while DeepSeek reaches 181.8 at its fastest setting. It does not track effort in any clear way, moving instead with how long the answer is. The answered rate holds above nine in ten from low to high among the models you can deploy, though it falls off at max for some, DeepSeek dropping to 85.0%.

Output speed (tokens/s)

Task duration (s)

Answered

GLM 5.3 FlashDeepSeek V4.1 FlashGPT-5.6 Luna

Memory required per model

The memory needed to keep the model itself resident when you host it yourself. The same model fits in less at lower precision, so the figures are shown per published build.

Memory required per model

The memory the weights need to stay resident. It does not change with reasoning effort, so there is one figure per model. Real use needs more than this: the working space grows with the length of the documents.

All figures

The answered figure is how often the model returned an answer; the pass rate is how often that answer was correct. On mobile, scroll the table horizontally.

ModelProviderReasoning effortCostPass rateAnsweredTrials (completed)
Single appTwo appsThree apps
GLM 5.3 FlashOllama CloudNone$0.025374.2%52.6%18.8%80.0%300
Low$0.001473.1%58.9%23.2%90.6%300
Medium$0.002177.4%75.8%35.7%91.5%300
High$0.001476.3%69.5%29.5%96.8%300
Max$0.002179.6%73.7%33.0%92.9%300
DeepSeek V4.1 FlashOllama CloudNone$0.001855.9%54.7%17.9%100.0%300
Low$0.003377.4%63.2%38.4%96.2%300
Medium$0.00478.5%63.2%30.4%99.3%300
High$0.003373.1%66.3%32.1%96.5%300
Max$0.013576.3%51.6%24.1%85.0%300
GPT-5.6 LunareferenceOpenAINone$0.050569.0%35.2%10.3%51.5%154
Low$0.01562.9%51.4%24.3%71.4%210
Medium$0.015358.2%51.3%17.9%74.7%221
High$0.013773.7%64.1%31.8%80.7%242
XHigh$0.013373.0%64.4%36.2%94.7%284
Max$0.014876.3%69.5%35.7%99.3%300

Cost is the model charge incurred by running one task. Development, maintenance and human review are excluded. The figure is the cost of a single task; where it falls below one cent it is shown with decimals. Running the model in your own environment removes this charge, leaving only the server cost.

How to choose

The same task can call for a different model depending on what you are optimising for. Below, common office jobs are split into the configuration we would recommend and the limit that remains even with it.

Getting through a high volume of routine work

Recommended: GLM 5.3 Flash at high effort

For work that stays inside one tool — pulling figures out of scanned documents into a list, say — the pass rate is 76.3%.

One task in four is still wrong. Where an error flows straight into the business — amounts, dates — design the process on the assumption that a person checks it.

Finishing work that crosses a spreadsheet and a document

Recommended: GLM 5.3 Flash at medium effort

For summarising figures from a sheet into a report document, the pass rate is 75.8% — the highest among the deployable models on work spanning two kinds of tool.

Dropping to low effort takes the pass rate down to 58.9%, the widest gap between adjacent settings we measured. Plan on keeping this one at medium.

Delegating work that crosses several tools

Recommended: DeepSeek V4.1 Flash at low effort

For work that reads a PDF, loads it into a spreadsheet and emails a manager, the pass rate is 38.4% — the highest among the deployable models on work spanning three kinds of tool.

No model gets much past four in ten; hand over ten and six need reworking. Budget for a person to check the result at the end.

Keeping it running while nobody is watching

Recommended: DeepSeek V4.1 Flash at medium effort

A 99.3% answered rate, the highest among the deployable models short of turning reasoning off entirely — for work where an unanswered trial is the thing you most want to avoid.

The answered rate is how often an answer came back at all, not how often it was right. Turning reasoning off does reach 100%, but the pass rate drops to 41.3%.

Cutting the time each task takes

Recommended: GLM 5.3 Flash at low effort

A median task takes 13.9 seconds — the shortest we measured apart from DeepSeek with reasoning off at 10.7 seconds — which helps when a batch of documents has to be cleared quickly.

Speed trades against accuracy: on work crossing several tools the pass rate falls to 23.2%. Duration and correctness do not improve together, so decide which one leads.

Reference: seeing where the ceiling sits

Recommended: GPT-5.6 Luna at max effort

A 99.3% answered rate, among the highest we measured, and a 35.7% pass rate even on work crossing several tools — about level with the best deployable model (38.4%).

It is not a model that can be placed in your own environment. It is listed to show how far the deployable models sit from that ceiling.

The right setup for your office work, with a purpose-built AI

What is above are general tendencies by kind of task. In real work the result shifts with how your documents are built and how your team checks what the model produces. We will work out the sums for your own work, and design a way of running it that holds up.

Evaluation conditions

Benchmark: OfficeBench, an academic benchmark whose paper and scoring code were published in 2024. The official 300 tasks and the official scoring programs are used unchanged. The corpus splits 93 / 95 / 112 tasks across one, two and three applications.

Models compared: GLM 5.3 Flash and DeepSeek V4.1 Flash (both on Ollama Cloud) are the models that can be deployed in your own environment. GPT-5.6 Luna (OpenAI) is evaluated only to give those two a scale, and cannot be placed in your own environment the same way. The deployable candidates were run at five reasoning efforts (None / Low / Medium / High / Max) and Luna at six, one more than the others. Every deployable effort completed all 300 trials.

Runtime: the two candidates were run on Ollama Cloud. The model behaves the same as it would in your own environment; speed depends on the compute you provide, so treat the speeds here as guidance when choosing a configuration.

Cost: each model's own published token rates multiplied by the tokens it actually used. The published rates are in USD, so no currency conversion is applied. Figures are the cost of a single task; as these fall below one cent they are shown to four decimal places. DeepSeek is priced at its base-window rate; its peak-window rate is not used. Running the model in your own environment removes this charge and leaves only the server cost; development, maintenance and human review are excluded either way.

Correctness: scoring mechanically checks the final file state — whether the requested value is present, whether the file is saved under the requested name, and whether unwanted values were left out. The pass rate is the share of completed trials that passed.

Speed: output tokens per second is output tokens divided by whole-call time, so it includes the wait before the first token and is not the model's intrinsic generation rate. Task duration is the median elapsed time for one task and includes our scoring step (about four seconds).

Failure attribution: a trial that produced no result is assigned to the provider, to our own harness, or to the model returning nothing. The first two are not model performance and are excluded from the answered figure.

Provenance: every figure is aggregated automatically from the run records (run-result.json). The aggregator and the tests that check its output are maintained separately from the page.

社内業務は専用AIで効率化しませんか?

扱う書類や確認手順が変われば、任せられる範囲も変わります。どの作業から始めるかの切り分けと、無理なく回る運用のかたちまで、まとめてご相談ください。