1. Workflow Overview
This document compares the transcription workflows of ITv2 and ITv3.
| Aspect | ITv2 | ITv3 |
|---|---|---|
| Workflow | 2 Annotators → 2 Reviewers → 1 Final Reviewer | 3 Annotators → 1 Reviewer |
| People per Task | 5 | 4 |
| Review Stages | 3 checking stages | 1 checking stage |
| Final Reviewer | Yes | No |
Key Difference
ITv2 uses more review stages and more people per task, providing greater review redundancy. ITv3 uses three annotators followed by a single reviewer, resulting in a simpler workflow with fewer people involved.
2. Comparison Overview
Both ITv2 and ITv3 contain 385 completed tasks from ume Batch 5.
| Aspect | ITv2 | ITv3 |
|---|---|---|
| Completed Tasks | 385 | 385 |
| Batch | ume Batch 5 | ume Batch 5 |
| Batch ID | ume_batch_5 |
ume_batch_5_v3 |
| Final State | Finalised | Reviewed |
| Images / Tasks | Different set | Different set |
Important Note
The ITv2 and ITv3 datasets contain different images and different tasks. The same 385 pages were not processed through both workflows.
Accuracy is calculated by comparing each annotator’s transcript with the finalised transcript of the same workflow.
The comparison covers accuracy, quality, efficiency, speed, and estimated cost, along with qualitative feedback from reviewers.
3. Comparison Table
| Category | Metric | ITv2 | ITv3 | Key Finding |
|---|---|---|---|---|
| Accuracy | Annotation Error | 3.77% | 4.87% | ITv2 lower |
| Review Error | 1.25% | N/A | ITv2 only | |
| Quality | Annotator Agreement | 0.939 | 0.922 | ITv2 slightly higher |
| Reviewer Agreement | 0.982 | N/A | ITv2 only | |
| Annotator vs. Reviewer | 3.46% | 4.87% | ITv2 lower | |
| Review Redundancy | 2 Reviewers + Final Reviewer | 1 Reviewer | ITv2 has more checks | |
| Efficiency | Tasks with Any Rejection | 6 (1.56%) | 12 (3.12%) | ITv3 about 2× higher |
| Annotator Times Sent Back | 5 | 15 | ITv2 lower | |
| Reviewer Times Sent Back | 8 | N/A | ITv2 only | |
| Total Rework | 13 | 15 | Almost the same | |
| Mean Rework per Task | 0.034 | 0.039 | Almost the same | |
| Total Works | 1,925 | 1,540 | ITv3 uses fewer people | |
| Speed | Median Annotator Hold | 20.5 min | 19.6 min | Nearly the same |
| Median Reviewer Hold | 6.4 min | 12.2 min | ITv3 reviewer hold is ~2× longer | |
| Median Final Reviewer Hold | 2.6 min | N/A | ITv2 only | |
| Median Total Hold per Task | 124.2 min | 186.9 min | ITv2 lower | |
| Cost | Total Cost | ₹246,858 | ₹247,199 | Almost the same |
| Average Cost per Task | ₹641.2 | ₹642.1 | Almost the same |
3.1 Key Findings
The comparison can be summarized across the major dimensions as follows:
| Area | What the Data Shows | Implication |
|---|---|---|
| Accuracy | ITv2 annotation error is 3.77% compared with 4.87% for ITv3. | ITv2 annotators are closer to their workflow’s finished output. |
| Quality | ITv2 annotator agreement is 0.939 compared with 0.922 for ITv3. | ITv2 shows slightly higher annotator consistency. |
| Reviewing | ITv2 has 2 reviewers + a final reviewer, while ITv3 has 1 reviewer. | ITv2 provides more review redundancy; ITv3 provides a simpler review process. |
| Rework | Total rework is 13 for ITv2 and 15 for ITv3. | Rework is very similar between the workflows. |
| People Required | ITv2 requires 5 people per task, while ITv3 requires 4. | ITv3 requires fewer people to complete a task. |
| Reviewer Experience | ITv3 allows the reviewer to compare three annotation versions together. | Reviewers prefer the simpler ITv3 workflow. |
| Speed | Annotator hold time is nearly identical, but ITv3 reviewer hold time is longer. | ITv3 has higher total hold time per task. |
| Cost | Average cost is ₹641.2 for ITv2 and ₹642.1 for ITv3. | Estimated cost is effectively the same. |
3.2 Calculations
| Section | Metric | ITv2 Formula / Definition | ITv3 Formula / Definition | Notes |
|---|---|---|---|---|
| Base Setup | Finished Page (F) | Finalised transcript | Review transcript | — |
| Accuracy | Annotation Error | (A1 vs F + A2 vs F) / 2 |
(A1 vs F + A2 vs F + A3 vs F) / 3 |
Mean across all 385 pages |
| Review Error | (R1 vs F + R2 vs F) / 2 |
N/A (reviewer = F) | Mean across all 385 pages | |
| Quality | Annotator Agreement | A1 vs A2 |
(A1 vs A2 + A2 vs A3 + A1 vs A3) / 3 |
Mean across all 385 pages |
| Reviewer Agreement | R1 vs R2 |
N/A | Mean across all 385 pages | |
| Annotator vs. Reviewer | (A1 vs R1 + A2 vs R1 + A1 vs R2 + A2 vs R2) / 4 |
(A1 vs R + A2 vs R + A3 vs R) / 3 |
Mean across all 385 pages | |
| Review Redundancy | 2 reviewers + 1 final reviewer (3 checkers) | 1 reviewer (1 checker) | Design fact | |
| Efficiency | Tasks with Any Rejection | (pages with at least 1 send-back) / 385 × 100 |
(pages with at least 1 send-back) / 385 × 100 |
Rejection rate across all pages |
| Annotator Times Sent Back | A1 send-backs + A2 send-backs | A1 send-backs + A2 send-backs + A3 send-backs | Total count of annotator send-backs | |
| Reviewer Times Sent Back | R1 send-backs + R2 send-backs | N/A | Total count of reviewer send-backs | |
| Total Rework | Annotator send-backs + reviewer send-backs | Annotator send-backs | Total send-backs combined | |
| Mean Rework per Task | Total rework / 385 | Total rework / 385 | Average do-overs per task | |
| Total Works | 385 × 5 |
385 × 4 |
Total person-jobs without rejections | |
| Speed | Individual Hold Time | (Submit time − Assign time) / 60 |
(Submit time − Assign time) / 60 |
Median per role |
| Total Hold per Task | A1 + A2 + R1 + R2 + FR | A1 + A2 + A3 + R | Median of all 385 task totals | |
| Cost Estimation | Total Cost (385 tasks) | Addition of total cost | Addition of total cost | Total cost added together |
| Average Cost | Total cost / 385 | Total cost / 385 | Cost per task |
3.3 Metric Definitions
Accuracy
Annotation Error:
Measures how different the annotator transcript is from that workflow’s finished transcript. A lower value means the annotator output is closer to the final output.
Review Error:
Measures how different the reviewer transcript is from the finished transcript. ITv3 is N/A because the reviewer produces the finished transcript.
Quality
Annotator Agreement:
Measures how similar annotators are to each other when working on the same image.
-
ITv2: 2 annotators
-
ITv3: 3 annotators
Reviewer Agreement:
Measures how similar ITv2’s two reviewers are to each other. ITv3 is N/A because it has only one reviewer.
Annotator vs. Reviewer:
Measures how different the annotator transcript is from the reviewer transcript. A lower value indicates that less change was made during review.
For ITv3, this value is the same as the annotation error because the reviewer’s transcript is also the finished transcript.
Review Redundancy:
Measures how many people check the work after annotation.
Efficiency
Tasks with Any Rejection:
Number of the 385 tasks that were rejected or sent back at least once.
Total Rework:
Total number of times a task was sent back and required additional work.
Mean Rework per Task:
Total rework divided by 385 completed tasks.
Annotator Times Sent Back:
Number of times an annotator was rejected and required to redo the task.
Reviewer Times Sent Back:
Number of times an ITv2 reviewer was sent back by the final reviewer.
Total Works:
Base number of person-jobs required to complete 385 tasks without additional rework.
-
ITv2: 385 × 5 = 1,925
-
ITv3: 385 × 4 = 1,540
Speed
Individual Hold Time:
Time between assignment and submission:
(Submit Time − Assign Time) / 60
The median is used for each role.
Total Hold per Task:
Sum of the hold times for all people involved in completing a task.
-
ITv2: A1 + A2 + R1 + R2 + Final Reviewer
-
ITv3: A1 + A2 + A3 + Reviewer
The median is then calculated across all 385 tasks.
4. Comparison Conclusion
| Dimension | Result |
|---|---|
| Accuracy | ITv2 performs better, with lower annotation error (3.77% vs 4.87%). |
| Quality | ITv2 has slightly higher annotator agreement (0.939 vs 0.922) and more review redundancy. |
| Review Structure | ITv2 has three checking stages, while ITv3 has one reviewer comparing three annotations. |
| Efficiency | ITv3 uses fewer people, while total rework remains almost the same. |
| Rejections | Rejections are low in both workflows, although ITv3 has a higher rejected-task rate (3.12% vs 1.56%). |
| Speed | Annotator time is nearly identical, but ITv3 has a longer reviewer hold time and higher total hold time per task. |
| Cost | Estimated costs are almost identical: ₹641.2 vs ₹642.1 per task. |
| Reviewer Preference | Reviewers prefer ITv3 because it provides a simpler workflow and allows three annotation versions to be compared together. |
Overall Conclusion
ITv2 provides stronger measured accuracy, slightly higher annotator agreement, and greater review redundancy. ITv3 provides a simpler workflow with fewer people and fewer review stages. Rework and estimated cost are nearly identical between the two workflows, while ITv3 has a higher reviewer and total hold time.
On these two separate 385-task samples, ITv2 shows stronger measured quality and review control, while ITv3 provides a simpler reviewer experience and requires fewer people.
5. Feedback from Reviewers on the Two Tools
Reviewer Preference
Reviewers generally prefer ITv3 because of its simpler workflow.
| Feedback Area | Reviewer Observation |
|---|---|
| Workflow Structure | 3 Annotators → 1 Reviewer is simpler |
| Comparison | The reviewer can compare three annotation versions together |
| Decision-Making | Three versions provide more choices when determining the correct transcription |
| Review Stages | There is no second reviewer or final reviewer |
| Reviewer Experience | The workflow feels simpler and more straightforward |
| Overall Preference | Reviewers prefer ITv3 |
Reviewer Feedback Summary
Reviewers prefer ITv3 because they can compare three annotation versions together and make their decision in a single review stage. They find this simpler than the multi-stage ITv2 workflow, which includes two reviewers and a final reviewer.
The main advantage identified by reviewers is that having three annotation versions available at the same time provides more choices and makes the comparison process clearer.
6. Overall Takeaway
| ITv2 Strengths | ITv3 Strengths |
|---|---|
| Lower annotation error | Simpler workflow |
| Slightly higher annotator agreement | Fewer people required |
| More review redundancy | Three annotations available to one reviewer |
| Lower reviewer hold time | No second reviewer or final reviewer |
| Lower total hold time | Preferred by reviewers |
| Slightly lower estimated cost | Similar overall rework |
Final Takeaway
ITv2 is stronger from a measured accuracy and review-control perspective, while ITv3 is stronger from a workflow simplicity and reviewer-experience perspective.
The data also shows that cost and rework are very similar between the two workflows, meaning the primary trade-off is between ITv2’s additional review redundancy and ITv3’s simpler review structure.
Below is the GitHub repository containing the data used for verification of the comparison analysis.