Established author-level metrics — the h-index, the Journal Impact Factor and journal quartile rankings — are precisely the instruments the San Francisco Declaration on Research Assessment (DORA) cautions against using to judge individuals. They are field-blind, so scholars in low-citation disciplines such as education appear systematically weaker; and they reward citation volume in high-prestige journals while remaining blind to what the Scholarship of Teaching and Learning (SoTL) actually produces. Version 1 of the SoTL Index answered this with a transparent, field-normalised alternative: five dimension scores, a weighted composite, and a 1–7 developmental spectrum. This paper describes version 2 — and the reason there will be no scores in it at all. On reflection, and turning the instrument's own critical commitments back on itself, we concluded that a composite score is a ranking technology regardless of how responsibly it is displayed, and that even a worded developmental band carries the same ordinal verdict in gentler clothing. Version 2 therefore severs the score at the computation layer: the instrument now describes a footprint of teaching-and-learning scholarship — counts, open-access shares stated as "X of Y", the actual venue names, active years, and field context — and pairs every observation with a stated blind spot. It refuses to render at all when the open record is too thin to describe fairly, discloses what it actively distorts rather than only what it misses, treats recognition as narrative rather than points, and names a contestation and custodianship mechanism. We set out the v1 design, the argument that dismantled it, the v2 instrument, a worked example, and the limits of describing as well as of scoring.
Research assessment is dominated by citation-based author metrics, chief among them the h-index (Hirsch, 2005). These metrics carry three well-documented problems for the Scholarship of Teaching and Learning. First, they are field-blind. Citation densities differ by orders of magnitude across disciplines; a citation rate that is exceptional in education is unremarkable in molecular biology. Comparing raw counts across fields — as the h-index implicitly does — systematically disadvantages teaching-focused and social-science scholars. Second, they privilege journal prestige. The Journal Impact Factor and SCImago/Scopus quartiles measure the journal, not the work, yet are routinely read onto individuals. Third, they are blind to the outputs SoTL most values — open educational resources, students-as-partners work, sustained programmatic contribution, and dissemination into practice rather than into highly cited journals.
The San Francisco Declaration on Research Assessment (DORA, 2012), the Leiden Manifesto (Hicks et al., 2015) and the Hong Kong Principles (Moher et al., 2020) together call for assessment that is field-normalised, transparent, holistic, and focused on the merits of the work rather than the venue. The SoTL Index is a deliberate, modest attempt to operationalise those principles for teaching scholarship. What follows records how far v1 got — and where its own principles caught up with it.
Version 1 computed, for a scholar (identified by ORCID) or an institution, five dimensions from open OpenAlex data, each scored 0–100: reach (mean citations per output, scaled by a field factor derived from the ratio of global to field median journal citedness, clamped to 0.5–3.0); openness (open-access share); SoTL focus (share of journal outputs in recognised teaching-and-learning venues); collaboration (distinct co-authors against a reference of 25); and sustained contribution (active years against span and a six-year maturity reference). A user-weightable composite mapped to a 1–7 developmental band (Emerging · Building · Developing · Establishing · Embedding · Leading · Transforming). The design discipline was displayed transparency: every component always visible, weights adjustable live, a spectrum in place of a precise score, and an explicit framing as a reflection tool rather than a ranking of people.
Four arguments — the first and fourth each sufficient on its own — led to the withdrawal of the scored design. They apply the instrument's own commitments — DORA's, and the author's critical work on how assessment infrastructures produce idealised, rankable subjects — to the instrument itself.
v1's discipline was all at the display layer: show every component, allow re-weighting, band rather than rank. But any stored, orderable scalar can be extracted and sorted downstream by whoever holds two of them. The moment two scholars' composites exist, a league table exists in potential; no interface convention prevents a reader, a committee or a spreadsheet from realising it. If the commitment is that the instrument must be incapable of ranking a person — rather than merely well-mannered about it — the scalar cannot exist. v2 therefore removes it from the computation: the instrument stores and prints no 0–100, and emits no normalised, weighted or field-adjusted quantity of any kind. The plain counts it does print can, as §6 concedes, be re-aggregated by others — but the instrument itself never authors the scalar.
The 1–7 spectrum was adopted in v1 precisely to resist fine-grained ranking. On reflection it fails the same test more quietly: "Emerging" to "Transforming" is an ordered scale with a named bottom rung, and a scholar placed on rung two has been found wanting in words instead of digits. Developmental language softens the verdict; it does not withdraw it. v2 replaces the ladder with descriptions that are deliberately non-comparable between people: what the record contains, where it appears, how open it is, over what span — facts that characterise without ordering.
v1's optional sixth dimension awarded points for fellowships, national and institutional awards, editorial roles, keynotes, grants and leadership. These credentials are real — and unequally distributed, along lines the recognition literature and the author's own work would predict: they accrue to those already visible, already networked, already recognised. Scoring them converts the sector's historical distribution of recognition into the scholar's personal deficit. v2 treats recognition in two non-scored ways: a self-authored narrative claim, shown as written and never counted; and an outward-pointing diagnostic that reads the corpus for signals that the sector's recognition systems — not the scholar — may have under-credited sustained, open teaching scholarship. The diagnostic appears on a fixed, disclosed trigger — the presence of recognised teaching-and-learning venues together with a majority-open record — not on a hidden judgement.
OpenAlex, like every bibliographic index, is strongest on DOI-indexed journal articles and weakest on practice reports, open educational resources, books and chapters, community-engaged work, and Indigenous scholarship — the long tail where much teaching scholarship lives. An instrument that scores over this data does not merely miss that work; it actively renders its absence as a low number, reporting the infrastructure's blindness as the scholar's thinness. No calibration fixes this, because the error is not in the constants but in what gets to count as signal. v2's response is structural: a footprint is described, with the blind spot stated beside every observation; and where the record is too thin to describe fairly (fewer than five resolvable outputs, fewer than three outputs with a resolvable journal source, or no resolvable primary field), the instrument refuses to render anything at all, stating that a sparse open record is a limit of what open bibliographic data has captured, never a statement about the amount or worth of the work.
v2 reads the same open data — OpenAlex (Priem, Piwowar & Orr, 2022), CC0, queried from the browser, nothing stored — and emits only descriptive facts, each paired with its stated blind spot and a reflective prompt:
In institutional mode, only the first three facets are read: co-authorship and time span are author-level facts and are not rendered for an institutional sample.
Around the facets sit four structural features. An always-shown blind-spot statement lists what the instrument cannot see (teaching itself; mentoring; curriculum; non-DOI work) at the head of every reading, not in a footnote — and the in-tool disclosure names what the instrument actively gets wrong, not only what it misses: it under-reads education and SoTL citations against research baselines, can mis-tag interdisciplinary or emerging venues as non-SoTL, and mistakes a paywall — a publisher or APC barrier — for a scholar's choice. An adequacy gate refuses to render thin records, as §3.4 describes. A recognition panel takes self-authored narrative, displayed verbatim and never scored, alongside the outward under-recognition diagnostic. And a contestation and custodianship notice states that any description can be challenged and corrected, that a wrong description is a defect in the instrument rather than a fact about the scholar, that the venue taxonomy is a living list held by an editorial custodian group whose members will be publicly named (in formation, to be convened with sector bodies), and that a scholar may ask for their own footprint to be re-described or withheld. The instrument never generates a First Nations–specific reading or recognition; that space is held custodian-led, outside it, by design.
For the author's own public record (ORCID 0000-0002-1428-2921), read live in July 2026, v2's reading, in summary: a reading of 36 public outputs, primarily in the Social Sciences; 21 of 36 open access; 13 of 17 journal articles in recognised teaching-and-learning venues — among them the Journal of University Teaching & Learning Practice, the International Journal for Students as Partners, the Australian Journal of Indigenous Education and the Pacific Journal of Technology Enhanced Learning; roughly 200 citations across the corpus, read against the citation norms of a low-citation field; outputs recorded in 12 of the 14 years from 2013 to 2026. v1 rendered this same record as a single worded band on its 1–7 spectrum. Nothing a reader legitimately needs was lost in the translation: the facts, the venues and the span are all still present. What was lost is only the verdict — which is the point.
The SoTL Index runs free and login-free at ntlsn.com/sotl-index.html, entirely in the browser. The accompanying Responsible Research Assessment hub (ntlsn.com/responsible-assessment.html) curates the wider DORA toolkit. This methodology is released under CC BY 4.0; reuse and adaptation are encouraged with attribution.
Boyer, E. L. (1990). Scholarship Reconsidered: Priorities of the Professoriate. Carnegie Foundation for the Advancement of Teaching.
DORA. (2012). San Francisco Declaration on Research Assessment. https://sfdora.org/read/
Glassick, C. E., Huber, M. T., & Maeroff, G. I. (1997). Scholarship Assessed: Evaluation of the Professoriate. Jossey-Bass.
Hicks, D., Wouters, P., Waltman, L., de Rijcke, S., & Rafols, I. (2015). Bibliometrics: The Leiden Manifesto for research metrics. Nature, 520(7548), 429–431. https://doi.org/10.1038/520429a
Hirsch, J. E. (2005). An index to quantify an individual's scientific research output. PNAS, 102(46), 16569–16572. https://doi.org/10.1073/pnas.0507655102
Hutchings, P., & Shulman, L. S. (1999). The scholarship of teaching: New elaborations, new developments. Change, 31(5), 10–15.
Moher, D., Bouter, L., Kleinert, S., et al. (2020). The Hong Kong Principles for assessing researchers. PLoS Biology, 18(7), e3000737. https://doi.org/10.1371/journal.pbio.3000737
Priem, J., Piwowar, H., & Orr, R. (2022). OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. arXiv:2205.01833.