ITv2 vs ITv3: Transcription Workflow Comparison Report

1. Workflow Overview

This document compares the transcription workflows of ITv2 and ITv3.

Aspect ITv2 ITv3
Workflow 2 Annotators → 2 Reviewers → 1 Final Reviewer 3 Annotators → 1 Reviewer
People per Task 5 4
Review Stages 3 checking stages 1 checking stage
Final Reviewer Yes No

Key Difference

ITv2 uses more review stages and more people per task, providing greater review redundancy. ITv3 uses three annotators followed by a single reviewer, resulting in a simpler workflow with fewer people involved.


2. Comparison Overview

Both ITv2 and ITv3 contain 385 completed tasks from ume Batch 5.

Aspect ITv2 ITv3
Completed Tasks 385 385
Batch ume Batch 5 ume Batch 5
Batch ID ume_batch_5 ume_batch_5_v3
Final State Finalised Reviewed
Images / Tasks Different set Different set

Important Note

The ITv2 and ITv3 datasets contain different images and different tasks. The same 385 pages were not processed through both workflows.

Accuracy is calculated by comparing each annotator’s transcript with the finalised transcript of the same workflow.

The comparison covers accuracy, quality, efficiency, speed, and estimated cost, along with qualitative feedback from reviewers.


3. Comparison Table

Category Metric ITv2 ITv3 Key Finding
Accuracy Annotation Error 3.77% 4.87% ITv2 lower
Review Error 1.25% N/A ITv2 only
Quality Annotator Agreement 0.939 0.922 ITv2 slightly higher
Reviewer Agreement 0.982 N/A ITv2 only
Annotator vs. Reviewer 3.46% 4.87% ITv2 lower
Review Redundancy 2 Reviewers + Final Reviewer 1 Reviewer ITv2 has more checks
Efficiency Tasks with Any Rejection 6 (1.56%) 12 (3.12%) ITv3 about 2× higher
Annotator Times Sent Back 5 15 ITv2 lower
Reviewer Times Sent Back 8 N/A ITv2 only
Total Rework 13 15 Almost the same
Mean Rework per Task 0.034 0.039 Almost the same
Total Works 1,925 1,540 ITv3 uses fewer people
Speed Median Annotator Hold 20.5 min 19.6 min Nearly the same
Median Reviewer Hold 6.4 min 12.2 min ITv3 reviewer hold is ~2× longer
Median Final Reviewer Hold 2.6 min N/A ITv2 only
Median Total Hold per Task 124.2 min 186.9 min ITv2 lower
Cost Total Cost ₹246,858 ₹247,199 Almost the same
Average Cost per Task ₹641.2 ₹642.1 Almost the same

3.1 Key Findings

The comparison can be summarized across the major dimensions as follows:

Area What the Data Shows Implication
Accuracy ITv2 annotation error is 3.77% compared with 4.87% for ITv3. ITv2 annotators are closer to their workflow’s finished output.
Quality ITv2 annotator agreement is 0.939 compared with 0.922 for ITv3. ITv2 shows slightly higher annotator consistency.
Reviewing ITv2 has 2 reviewers + a final reviewer, while ITv3 has 1 reviewer. ITv2 provides more review redundancy; ITv3 provides a simpler review process.
Rework Total rework is 13 for ITv2 and 15 for ITv3. Rework is very similar between the workflows.
People Required ITv2 requires 5 people per task, while ITv3 requires 4. ITv3 requires fewer people to complete a task.
Reviewer Experience ITv3 allows the reviewer to compare three annotation versions together. Reviewers prefer the simpler ITv3 workflow.
Speed Annotator hold time is nearly identical, but ITv3 reviewer hold time is longer. ITv3 has higher total hold time per task.
Cost Average cost is ₹641.2 for ITv2 and ₹642.1 for ITv3. Estimated cost is effectively the same.

3.2 Calculations

Section Metric ITv2 Formula / Definition ITv3 Formula / Definition Notes
Base Setup Finished Page (F) Finalised transcript Review transcript
Accuracy Annotation Error (A1 vs F + A2 vs F) / 2 (A1 vs F + A2 vs F + A3 vs F) / 3 Mean across all 385 pages
Review Error (R1 vs F + R2 vs F) / 2 N/A (reviewer = F) Mean across all 385 pages
Quality Annotator Agreement A1 vs A2 (A1 vs A2 + A2 vs A3 + A1 vs A3) / 3 Mean across all 385 pages
Reviewer Agreement R1 vs R2 N/A Mean across all 385 pages
Annotator vs. Reviewer (A1 vs R1 + A2 vs R1 + A1 vs R2 + A2 vs R2) / 4 (A1 vs R + A2 vs R + A3 vs R) / 3 Mean across all 385 pages
Review Redundancy 2 reviewers + 1 final reviewer (3 checkers) 1 reviewer (1 checker) Design fact
Efficiency Tasks with Any Rejection (pages with at least 1 send-back) / 385 × 100 (pages with at least 1 send-back) / 385 × 100 Rejection rate across all pages
Annotator Times Sent Back A1 send-backs + A2 send-backs A1 send-backs + A2 send-backs + A3 send-backs Total count of annotator send-backs
Reviewer Times Sent Back R1 send-backs + R2 send-backs N/A Total count of reviewer send-backs
Total Rework Annotator send-backs + reviewer send-backs Annotator send-backs Total send-backs combined
Mean Rework per Task Total rework / 385 Total rework / 385 Average do-overs per task
Total Works 385 × 5 385 × 4 Total person-jobs without rejections
Speed Individual Hold Time (Submit time − Assign time) / 60 (Submit time − Assign time) / 60 Median per role
Total Hold per Task A1 + A2 + R1 + R2 + FR A1 + A2 + A3 + R Median of all 385 task totals
Cost Estimation Total Cost (385 tasks) Addition of total cost Addition of total cost Total cost added together
Average Cost Total cost / 385 Total cost / 385 Cost per task

3.3 Metric Definitions

Accuracy

Annotation Error:
Measures how different the annotator transcript is from that workflow’s finished transcript. A lower value means the annotator output is closer to the final output.

Review Error:
Measures how different the reviewer transcript is from the finished transcript. ITv3 is N/A because the reviewer produces the finished transcript.


Quality

Annotator Agreement:
Measures how similar annotators are to each other when working on the same image.

  • ITv2: 2 annotators

  • ITv3: 3 annotators

Reviewer Agreement:
Measures how similar ITv2’s two reviewers are to each other. ITv3 is N/A because it has only one reviewer.

Annotator vs. Reviewer:
Measures how different the annotator transcript is from the reviewer transcript. A lower value indicates that less change was made during review.

For ITv3, this value is the same as the annotation error because the reviewer’s transcript is also the finished transcript.

Review Redundancy:
Measures how many people check the work after annotation.


Efficiency

Tasks with Any Rejection:
Number of the 385 tasks that were rejected or sent back at least once.

Total Rework:
Total number of times a task was sent back and required additional work.

Mean Rework per Task:
Total rework divided by 385 completed tasks.

Annotator Times Sent Back:
Number of times an annotator was rejected and required to redo the task.

Reviewer Times Sent Back:
Number of times an ITv2 reviewer was sent back by the final reviewer.

Total Works:
Base number of person-jobs required to complete 385 tasks without additional rework.

  • ITv2: 385 × 5 = 1,925

  • ITv3: 385 × 4 = 1,540


Speed

Individual Hold Time:
Time between assignment and submission:

(Submit Time − Assign Time) / 60

The median is used for each role.

Total Hold per Task:
Sum of the hold times for all people involved in completing a task.

  • ITv2: A1 + A2 + R1 + R2 + Final Reviewer

  • ITv3: A1 + A2 + A3 + Reviewer

The median is then calculated across all 385 tasks.


4. Comparison Conclusion

Dimension Result
Accuracy ITv2 performs better, with lower annotation error (3.77% vs 4.87%).
Quality ITv2 has slightly higher annotator agreement (0.939 vs 0.922) and more review redundancy.
Review Structure ITv2 has three checking stages, while ITv3 has one reviewer comparing three annotations.
Efficiency ITv3 uses fewer people, while total rework remains almost the same.
Rejections Rejections are low in both workflows, although ITv3 has a higher rejected-task rate (3.12% vs 1.56%).
Speed Annotator time is nearly identical, but ITv3 has a longer reviewer hold time and higher total hold time per task.
Cost Estimated costs are almost identical: ₹641.2 vs ₹642.1 per task.
Reviewer Preference Reviewers prefer ITv3 because it provides a simpler workflow and allows three annotation versions to be compared together.

Overall Conclusion

ITv2 provides stronger measured accuracy, slightly higher annotator agreement, and greater review redundancy. ITv3 provides a simpler workflow with fewer people and fewer review stages. Rework and estimated cost are nearly identical between the two workflows, while ITv3 has a higher reviewer and total hold time.

On these two separate 385-task samples, ITv2 shows stronger measured quality and review control, while ITv3 provides a simpler reviewer experience and requires fewer people.


5. Feedback from Reviewers on the Two Tools

Reviewer Preference

Reviewers generally prefer ITv3 because of its simpler workflow.

Feedback Area Reviewer Observation
Workflow Structure 3 Annotators → 1 Reviewer is simpler
Comparison The reviewer can compare three annotation versions together
Decision-Making Three versions provide more choices when determining the correct transcription
Review Stages There is no second reviewer or final reviewer
Reviewer Experience The workflow feels simpler and more straightforward
Overall Preference Reviewers prefer ITv3

Reviewer Feedback Summary

Reviewers prefer ITv3 because they can compare three annotation versions together and make their decision in a single review stage. They find this simpler than the multi-stage ITv2 workflow, which includes two reviewers and a final reviewer.

The main advantage identified by reviewers is that having three annotation versions available at the same time provides more choices and makes the comparison process clearer.


6. Overall Takeaway

ITv2 Strengths ITv3 Strengths
Lower annotation error Simpler workflow
Slightly higher annotator agreement Fewer people required
More review redundancy Three annotations available to one reviewer
Lower reviewer hold time No second reviewer or final reviewer
Lower total hold time Preferred by reviewers
Slightly lower estimated cost Similar overall rework

Final Takeaway

ITv2 is stronger from a measured accuracy and review-control perspective, while ITv3 is stronger from a workflow simplicity and reviewer-experience perspective.

The data also shows that cost and rework are very similar between the two workflows, meaning the primary trade-off is between ITv2’s additional review redundancy and ITv3’s simpler review structure.

Below is the GitHub repository containing the data used for verification of the comparison analysis.