What Our Spot Checks Tell Us About Text-Boundary Annotation Quality

Our annotation task involves more than identifying where one Tibetan work ends and another begins. Annotators also identify structural material such as front matter, tables of contents, catalogs, and back matter.

To evaluate annotation quality, we analyzed the results of our final review process. The dataset contains 160 reviewer-detected issues across 64 volumes, reviewed between 11 June and 11 August 2026.

Our main questions were:

  • What kinds of errors were annotators making?

  • Did those errors change as annotators received repeated feedback from the final reviewer?

What kinds of errors did reviewers find?

Distribution of reviewer-detected errors

Most errors were related directly to identifying text boundaries.

Error type Errors Share
Under-segmentation 87 54.4%
Over-segmentation 35 21.9%
Wrong boundary 24 15.0%
TOC/catalog 8 5.0%
Front/back matter 3 1.9%
Title issue 2 1.3%
Incomplete segment 1 0.6%

Together, under-segmentation, over-segmentation, and wrong-boundary errors account for 91.3% of all reviewer findings.

The largest problem was under-segmentation: annotators sometimes treated multiple works as a single segment and missed boundaries that should have been added.

Structural errors involving TOCs, catalogs, front matter, and back matter were comparatively rare.

How did the errors change over time?

Looking only at the total number of errors hides an important part of the story. Different types of errors changed in different ways.

Trend of reviewer-detected errors

We divided the review period into three stages:

Period Affected volumes Total errors Errors per affected volume
Jun 11–30 24 58 2.42
Jul 1–16 32 89 2.78
Jul 17–Aug 11 8 13 1.63

The data does not show a simple continuous decline. Errors increased during the first half of July, partly because several particularly difficult volumes produced large clusters of mistakes.

After mid-July, however, the error profile changes substantially.

Under-segmentation

Under-segmentation was the most persistent problem.

Period Errors Errors per affected volume
Jun 11–30 33 1.38
Jul 1–16 50 1.56
Jul 17–Aug 11 4 0.50

There was little improvement initially, but after mid-July the rate dropped by about 68%, from 1.56 to 0.50 errors per affected volume.

This suggests that annotators became substantially better at recognizing when additional works still needed to be separated.

Over-segmentation

Over-segmentation showed the strongest change.

Period Errors
Jun 11–30 8
Jul 1–16 27
Jul 17–Aug 11 0

Over-segmentation became particularly visible during early and mid-July, although some of the increase came from a few unusually difficult volumes.

More importantly, no over-segmentation errors were recorded after 16 July.

This suggests that repeated feedback may have successfully corrected the tendency to split works too aggressively.

Wrong boundary placement

Wrong-boundary errors behaved differently.

Period Errors Errors per affected volume
Jun 11–30 9 0.38
Jul 1–16 9 0.28
Jul 17–Aug 11 6 0.75

This category did not improve in the same way.

By the final period, wrong-boundary errors accounted for 46.2% of all remaining issues.

This may indicate a progression in annotation difficulty. Earlier errors were often about whether a work needed to be split at all. Later, annotators were more often identifying the existence of a boundary correctly but struggling with exactly where that boundary should be placed.

This is likely the most important area for future reviewer feedback.

TOC, catalog, front matter, and back matter

Structural classification errors were relatively uncommon throughout the project.

TOC/catalog errors followed a pattern of:

5 → 1 → 2

Front/back-matter errors followed:

2 → 1 → 0

No front- or back-matter classification errors were recorded after 12 July.

Because the total number of these errors is small, we should avoid drawing strong statistical conclusions, but the results suggest that annotators were generally successful at distinguishing structural material from actual works.

What does this tell us about the feedback process?

The data suggests that annotation quality did improve, but not simply because every type of error steadily decreased.

Instead, the nature of the errors changed.

Earlier in the process, reviewers frequently found fundamental segmentation mistakes:

“Should this be one work or several?”

By the later review period, under-segmentation had fallen sharply and over-segmentation had disappeared. The remaining problems were increasingly about the more subtle question:

“We know there is a boundary, but exactly where should it be?”

This is an encouraging pattern because it suggests that annotators were learning from reviewer feedback and moving from broad segmentation mistakes toward more precise and difficult boundary decisions.

One important limitation

The current datasets contain only volumes where the reviewer recorded at least one issue.

They do not include reviewed volumes that passed with zero errors.

Because of this, we cannot calculate the true overall error rate or say precisely how much the pass rate improved over time.

For future review rounds, recording the total number of volumes reviewed—including zero-error volumes—would allow us to measure stronger metrics such as:

percentage of reviewed volumes passing without correction
or
errors per 100 reviewed text boundaries.

Even with this limitation, the current data provides a useful signal: after repeated reviewer feedback, major segmentation errors became less common, some error categories disappeared, and the remaining corrections became increasingly focused on precise boundary placement.