Appendix F — Change Log

This appendix records the exact changes applied in the v1.1 release relative to the as-submitted defended dissertation, followed by the additional corrections made in the subsequent v1.1.1 point release. These corrections are reflected in the current HTML and PDF renders generated from the same source. The list below is intended to be complete for the changes made in these passes. No new data were collected and no new analyses were introduced.

F.1 Change Register

Corrections are grouped by release below. The key interpretive and procedural changes are described more fully in the correction notes that follow this register.

F.1.1 v1.1

Location Issue in defended release Correction Reason
Literature Review footnote Unresolved *THEORY SECTION* placeholder Replaced with a direct cross-reference to the later theory section (@sec-theory) Remove draft residue and point readers to the actual theory discussion
Methods, retention hypotheses Leaked TODO rendered in public output because the hiding wrapper was misspelled as contents-hidden Removed the leaked note and rewrote the subsection as regular prose Remove visible draft residue and make the subsection publication-safe
Methods, retention hypotheses Retention was stated in Methods as H3a: Change in OEE varies with treatment, while defended Results evaluated retention using changes in task completion time and uncorrected error count Updated the Methods block so retention is stated and tested in terms of TCT and UCE, matching the defended Results chapter Reconcile Methods with the already-defended analysis without introducing a new retention analysis
Methods, legacy analysis-plan wording A few remaining Methods passages still reflected the original analysis plan, including references to OEE in retention context and hardware-specific treatment labels Added footnotes clarifying the final defended analysis path and the equivalence of HMDAR/HMDMR with the simpler AR/MR labels used later in the manuscript Improve transparency around residual legacy wording without rewriting the defended prose
Results, H2b Q-Q plot Placeholder subtitle subtitle... Replaced with PWI and PAR show the clearest departures from normality Remove placeholder text and summarize the displayed pattern
Results, H2b bootstrap interpretation Hidden TODO: FIX THIS remained in source and the visible typo more than than remained in prose Removed the hidden note, corrected the typo to more than, and softened the sentence to report what the estimates suggest rather than overstate certainty Editorial cleanup and more proportionate interpretation
Results, H1b model-comparison table compare_performance() emitted stdout into the rendered body text before the table Captured the emitted stdout so it no longer renders; the comparison table itself is unchanged Remove render artifact without changing the analysis
Results, H1b model-comparison table Candidate model 4 showed blank conditional \(R^2\) and ICC values without explanation Added an explicit note that candidate model 4 produced a singular fit, so conditional \(R^2\) and ICC were unavailable Make the omitted values transparent to readers
Results, sample and exclusion notes The scope of the participant 1063 exclusion, the recall analysis sample, and the H2b-specific outlier removal were implicit in the defended text Added footnotes clarifying the 1063 exclusion, the recall sample composition, and the fact that the H2b outlier removal affected only that analysis; corrected a nearby typographical error in the H2b outlier paragraph Improve transparency around analysis-specific sample definitions without revising the defended narrative
Results, reporting corrections A demographics sentence reused the Lego p-value for Education; one H3a sentence mismatched iUCE with an iTCT figure; one H3b figure caption used Uncounted instead of Uncorrected; and one typo read revering Corrected the Education p-value reference, aligned the H3a sentence with the iTCT figure, standardized the H3b caption to Uncorrected, and fixed the typo to reversing Correct concrete reporting and labeling errors without changing the underlying analyses
Results and Conclusions, H3b interpretation The defended text interprets the selected additive ZINB retention model as if gap effects were additive and treatment-specific in slope Added footnotes in Results and Conclusions that point to Section F.2.1, where the interpretive correction and its implications are discussed without rewriting the defended narrative in place Make a substantive interpretive correction transparent while preserving the defended text
Landing page, HTML edition Public HTML opened directly to the abstract, without a visible dissertation identity block or committee listing Added an HTML-only dissertation identification block before the abstract, including title, author, degree context, date, and approved committee list Make the HTML edition legible as a dissertation when opened directly
Landing page, HTML edition Public HTML did not visibly distinguish the corrected v1.1 release from the as-defended dissertation or provide a direct path to the project source Added an HTML-only v1.1 callout on the landing page with direct links to Appendix F, the as-defended dissertation PDF on GitHub, and the project repository on GitHub Make release status, provenance, and source access clear at the point of entry
Appendix D, Data Organization Visible TODO lines in public appendix text Removed the TODO lines Remove visible draft residue
Appendix E, Institutional Review Board Approval Appendix lead-in did not point readers to the defended public artifact Replaced the lead-in with a direct link to the standalone as-submitted IRB packet on GitHub Improve traceability to the defended IRB materials

F.1.2 v1.1.1

Location Issue in defended release Correction Reason
Results \(H_{2b}\) and Conclusions (recall) The recall reliance comparison is stated with its direction reversed, describing the augmented conditions as reducing paper reliance relative to PWI, accompanied by two misprinted magnitudes (0.686, 0.612) Added a correction note (Section F.2.2) and footnotes establishing that PWI and PAR showed the lowest paper reliance at recall and the head-mounted AR and MR the highest; Table 5.18 is correct and unchanged Corrects a directional reporting error that propagated into the recall interpretation and conclusions, without altering the defended table or the \(H_{2b}\) determination
Results \(H_{2a}\) and \(H_{2b}\) bootstrap The complementary bootstrap was described as five replications averaged together, with the across-replication dispersion reported as evidence of stability Added a correction note (Section F.2.3) and a footnote explaining that a seeding error made the five replications identical, so the reported dispersion is structurally zero and the estimates are those of a single bootstrap Corrects a misleading procedural description; point estimates, intervals, and the \(H_{2a}\) and \(H_{2b}\) conclusions are unchanged
Conclusions, learning results-summary table (\(H_{1b}\) row) The \(H_{1b}\) row states the model controlled for “iTCT” (the retention-phase increase in TCT); the defended model controlled for initial TCT (\(TCT_0\)) Added a footnote at Table 6.1 clarifying that the covariate should be read as the initial TCT (\(TCT_0\)); the defended table is left as displayed Clarifies the covariate label transparently without altering the defended summary table or the \(H_{1b}\) determination
Conclusions, Empirical Contributions The Empirical Contributions summary groups PAR with PWI as showing higher error rates, contradicting the chapter’s own \(H_{1c}\) result and summary table (PAR is statistically low-error, indistinguishable from AR and MR) Added a footnote at the affected sentence noting that PAR was significantly lower in errors than PWI and groups with the augmented conditions on accuracy. The abstract carries the same augmented-vs-traditional framing on the speed axis (“augmented technologies generally led to … slower initial performance”), where PAR is the analogous fast exception; that statement is reasonably hedged and is left as written Corrects an internal inconsistency without altering the defended sentence; the \(H_{1c}\) determination is unchanged
Results, \(H_{2a}\) OEE bootstrap discussion The PWI-MR average OEE difference is misprinted as 0.307; the correct value (the mean of the absolute Mean/Median/HL differences) is 0.614, exactly twice the printed figure Added a footnote flagging the misprint and the corrected value (0.614); the PWI-PAR (0.687) and PWI-AR (0.585) figures are correct, and once corrected the three PWI-versus-augmented gaps (0.687, 0.585, 0.614) are comparable rather than MR appearing smallest Corrects a misprinted descriptive statistic; no augmented-vs-augmented difference was significant, so the \(H_{2a}\) determination is unchanged
Results, \(H_{3b}\) model table (Table 5.24) The “Adj Coef” column halves every coefficient to undo the 2x response scaling, but the treatment coefficients are scale-invariant rate ratios that should not be halved; the PWI-vs-AR factor is stated as 0.197 instead of 0.394 Extended the H3b correction note (Section F.2.1) to explain that the halving applies only to the intercept and that the PWI factor is 0.394 Corrects a coefficient-rescaling error within the already-noted H3b interpretation; sign, significance, and the \(H_{3b}\) conclusion are unchanged
Results, \(H_{2a}\) OEE bootstrap introduction States that the bootstrap “increases the sample size and statistical power”; resampling does neither, since it reuses the observed data Extended the bootstrap correction note (Section F.2.3) to clarify that resampling improves only the precision of the sampling-distribution estimate, not the effective sample size or power Corrects a methodological misstatement; the bootstrap is a complementary check and the \(H_{2a}\) conclusion is unchanged
Various (Results) Low-severity, conclusion-neutral slips: the Poisson model-comparison BIC claim (F-008; inert, both Poisson models are discarded for overdispersion); the iTCT median ordering (F-009; PAR highest, no treatment effect); the H1b fixed-vs-random correlation gloss (F-001; -0.75 vs -0.85, same sign and magnitude); the H1c “IQR and Z-score” wording (F-025; IQR only, a robustness check); the TLX Q2 boundary figures (G-004; off by hundredths, qualitative point holds); and a duplicated SUS figure caption (G-009; readable from its own title). Two bibliography-metadata items were corrected in a Zotero bibliography pass: the four cited skeletal @article stubs that rendered “(n.d.)” (F-034) were retyped as @misc web resources with organization authors, source URLs, and 2022 access dates (publication year left n.d., as these are undated web resources); and the malformed year in the uncited konig2018embod entry (F-035) was corrected to 2018. Flagged by footnote (F-008, F-009, F-001) or noted here (F-025, G-004, G-009); no defended number or table is altered; F-034 updates only the four affected bibliography entries (organization author, access date, URL) and F-035 is non-rendering General cleanup of minor reporting and descriptive slips; none affects a hypothesis decision, estimate, test, or conclusion
Various (Results, Conclusions, Methods, Problem Statement, Literature Review, change log) Residual copy and render defects: repeated words, eponym and term misspellings (e.g. “Kruskall-”/“Kurskal-Wallis”, “Heart” for Hart), verb-form slips, a “left-skewed” label that should read “right-skewed”, an orphan “RQ1” label, the affordance name “Egocentric Display” (should be “User-Centric Display”), “System Usability Score” (should be “Scale”), a figure cross-reference and a citation missing their leading “@”, a spurious squared superscript on a negative rank-biserial correlation, and “Appendix Appendix F” double-prints Corrected each inline in the rendered build Editorial and render cleanup only; no statistic, value, or claim is affected

F.2 Correction Notes

These notes record the substantive interpretive and procedural corrections referenced in the change register, each flagged at the affected passages by footnote rather than by rewriting the defended text.

F.2.1 v1.1 H3b Interpretation Correction Note

The defended \(H_{3b}\) discussion interprets the selected additive zero-inflated negative binomial model as if the effect of gap were additive and treatment-specific in slope. That is not correct. In the selected model, count-model effects are multiplicative on the log-link scale, and no gap-by-treatment interaction term is included. Accordingly, the model supports the conclusion that expected increases in UCE differ by treatment level across the observed retention intervals, with PWI lower than the reference AR condition. It does not support a claim that PWI changes the per-day rate of increase relative to AR, nor that treatment-specific decay slopes were estimated by the selected model.

This does not eliminate the \(H_{3b}\) signal, but it narrows its interpretation. The defended phrasing about a lower rate of error increase should therefore be read as overstated. A more accurate reading is that the fitted model provides preliminary evidence of differences in expected retention-phase error increase by treatment, while leaving the mechanism of that difference, including whether treatments differ in slope over time, unresolved. To preserve the defended narrative, v1.1 records that interpretive correction here and flags the affected passages in Results and Conclusions by footnote rather than rewriting them in place.

A related correction, added in the v1.1.1 pass: the “Adj Coef” column in Table 5.24 halves every coefficient to undo the doubling that was applied to the response. That halving is right for the intercept, which is an expected count and so scales with the response, but not for the treatment coefficients. Each of those is a multiplicative rate ratio, and a rate ratio does not change when the response is rescaled, because the factor of two cancels between the two groups being compared; the coefficients should therefore be left as estimated, not halved. Read correctly, PWI’s expected error increase is 0.394 of the reference AR condition’s, not the halved 0.197: a shift in the overall level by treatment, consistent with the reading above, rather than a treatment-specific change in slope.

F.2.2 v1.1.1 H2b Reliance Direction Correction Note

The recall-phase reliance discussion (\(H_{2b}\)) describes the augmented conditions as reducing reliance on the paper work instructions relative to PWI. That direction is reversed. As reported, the Results summary states that “AR and MR reduce reliance by 0.686 and 0.612 seconds per reference … relative to the PWI instructional method” and that “AR and MR … reduce reliance more than PAR and PWI,” and the Conclusions restate this as “reduced reliance on instructions for AR and MR users” (recall results-summary table; recall interpretation; Collective Insights).

The underlying bootstrap table (Table 5.18) is correct; the prose reads it backward. The pairwise differences are signed as PWI minus treatment and are negative for the augmented conditions (the mean difference is about -0.99 s for PWI versus AR and -0.88 s for PWI versus MR; the median and Hodges-Lehmann contrasts agree), so PWI sits at the low end of paper reliance, not the high end. Recomputed group reliance during recall confirms this on every measure (average references per task, total reference time per task, and time per reference), with PWI and PAR lowest and the head-mounted AR and MR conditions highest. The simplest measure agrees: the share of participants who never consulted paper at recall was PWI 57%, PAR 50%, AR 31%, MR 29%, so the augmented (particularly head-mounted) learners were the more likely to reference the instructions. The printed values 0.686 and 0.612 do not appear in Table 5.18 and are not the reliance differences.

This does not change the \(H_{2b}\) determination (reliance does vary with instructional method), and the Results caution that the specific pairwise differences are “not clear enough to declare” continues to govern. Reliance was low and heavily zero-weighted (143 of 212 recall tasks, and 22 of 53 participants, recorded no paper reference at all), so the comparison rests on a small number of referencers at modest per-group sample sizes. But the direction of the comparison was stated backward, and the interpretation it supports is affected: the recall reliance evidence does not indicate greater independence for the augmented conditions and, to the extent it points anywhere, indicates greater paper reliance at recall for the head-mounted AR and MR conditions. The defended learning-quality (\(H_{1c}\)) and recall-performance (\(H_{2a}\)) results do not depend on this measure and are unchanged. Per the errata-not-revision policy, the defended text and the (correct) table are left in place; this note records the correction and the affected passages are flagged by footnote.

Affected passages: in Results, the \(H_{2b}\) bootstrap summary and Result; in Conclusions, the recall results-summary table (\(H_{2b}\) row), the recall interpretation, and Collective Insights.

F.2.3 v1.1.1 Bootstrap Replication Seeding Correction Note

The H2a (OEE) and H2b (PWI reliance) complementary bootstrap analyses were described as five replications of 10,000 resamples each, averaged across replications, with the across-replication dispersion of bias and standard error reported as evidence of stability (the sd_Bias and sd_SE columns in Table 5.16 and Table 5.18). A seeding error made those five replications identical: the random seed was reset to the same fixed value immediately before each resampling call, so every replication drew the same resamples. Consequently, the averaging across replications was a no-op, and the reported across-replication dispersion is structurally zero rather than a measured quantity. The “five replications … averaged” description and the appeal to low across-replicate dispersion as evidence of reliability are therefore not supported as written.

This does not change the estimates or the conclusions. The tabulated differences, confidence intervals, biases, and standard errors are exactly those of a single valid 10,000-resample percentile bootstrap, which is itself a legitimate analysis; only the replication framing and the zero-valued dispersion columns are affected. The reported pairwise differences are deterministic functions of the full sample and do not depend on the seed. The primary inferences for both hypotheses rest on the Kruskal-Wallis tests, which are independent of this bootstrap, and the bootstrap is reported only as a complementary robustness check. The H2a and H2b results and their interpretations therefore stand as defended.

A corrected implementation, using a single high-resolution bootstrap without the per-replication reseed, is carried in follow-on work derived from this dissertation; it reproduces the same pattern of pairwise differences and leaves the substantive findings unchanged.

A related correction, added in the v1.1.1 pass: the introduction to this analysis states that the replication process “effectively increases the sample size and statistical power of the analysis.” Resampling does neither. The bootstrap reuses the observed data, so it adds no information and cannot raise the effective sample size or the power to detect an effect; both are fixed by the number of participants and the size of the observed effect. Additional resamples improve only the precision with which the sampling distribution and its confidence intervals are estimated (lower Monte Carlo error), not the power of the analysis.

F.3 Interpretation Boundaries

  • No new participants, data, or analyses were added for v1.1 or v1.1.1.
  • The most substantive corrections are the Methods/Results retention reconciliation (v1.1) and the corrected direction of the recall reliance comparison (v1.1.1; see the H2b reliance note above).
  • Other changes in these passes are editorial, render-related, or transparency-related.

F.4 Intentional Deferrals

The v1.1 and v1.1.1 releases are correction, reconciliation, and publication-safety passes, not a substantive revision of the defended dissertation. Several issues identified during review were therefore left unchanged by design. These include:

  • broader rhetorical tightening of H2b and retention claims beyond the specific corrections documented above
  • new sensitivity analyses or robustness checks for exclusions, outliers, or model choices
  • broader revision of the qualitative-analysis method or its presentation
  • deeper softening of appendix material beyond visible public-facing draft residue and link/provenance fixes
  • a full audit of all inline statistics beyond the specific exposed reporting errors corrected in these passes

These items were judged to fall outside the intended scope of an errata-style release. Where appropriate, they are better addressed in future paper manuscripts, follow-on analysis, or later revision work rather than in the corrected dissertation release itself.

F.5 Repository Diff

A repository compare view between the defended baseline tag and the v1.1 release tag is available on GitHub:

The additional v1.1.1 corrections relative to v1.1: