Technical Implementation Notes

From The Library Index Wiki

This page describes some of the technical issues we faced when building this website and certain decisions we made with its implementation.

Architecture

This website follows a design pattern pioneered by David Macfarlane for The Bibliographical Society:

  • keep the data in a normalised relational database, to take advantage of the non-redundant data storage, uniform structure, relational integrity, and content control
  • display the information using MediaWiki software, to take advantage of many ease-of-use features such as formatting, templates, and search
  • transform the data into the wiki using a site-specific program that builds the wiki pages from the database and updates them on the wiki when they change

The proof-of-concept site for this architecture was the London Book Trades system, where the database was originally built by Michael Turner but the presentation was difficult to navigate or understand.

Database

The database normalises the author or authors, type of article, article text, subjects, and citations to simplify the generation of not only the article pages, but wiki pages for subjects, authors, citations, etc.

It runs on a MySQL DBMS and contains 9,426 articles falling into various types based on article meta-data:

Count Article Type
3248 Reviews
2264 Articles
720 Bibliographical Notes
689 Recent Books and Periodicals
361 Correspondence
2144 All others

Note that the above article types have been coalesced to group "Review" with "Reviews", etc.

Several passes of custom data clean-up programs have been run on the data but in a non-destructive way. For example, the original article type is stored in a column named {article_type} and the cleaned up name in a separate column named {article_type_clean}.

Word Clouds

The word clouds are generated using a custom python program calling wordcloud. After some experimentation, we judged that 35 words created enough interest without overcrowding the image.

The words are chosen following these steps for each article:

  1. load the article text layer
  2. tokenise it (lowercase, alphabetic, stopword-filtered using English and non-English common short words)
  3. compute TF-IDF weights against the corpus of all words used in all articles
  4. collapse regular plurals onto their singulars where both would make the cloud ("variants" into "variant"), backfilling the freed slots,
  5. render a PNG word cloud of the top-35 words

TF-IDF (term frequency - inverse document frequency) is a statistical technique that measures how important a word is in a document compared to how often that word is used in all documents in a collection. TF is the proportion of words in the article being examined that are the target word; IDF is the proportion of the documents that use that word (it's actually the log of the proportion but we're trying to not get too mathematical here). Words that are used often in one article but infrequently in other articles get higher scores. The wordcloud uses the top 35 words ranked by TF-IDF score.

We also place the 10 top scoring words in an invisible comment in the page so they can be found by the search engine.

Details Table

There are two author name slots in the Details table in order to show how the author or authors were credited in the original publication and to give the normalised form of the author's name that we used to build author pages for each identified author. You can click on any of the names in the first Authors box to see what they wrote for The Library.

Subjects and Subject Categories

Subjects in the catalog are assigned by a small program that works from a controlled vocabulary drawn up by hand. For each subject, the vocabulary lists a set of seed terms, such as characteristic words and names, and allows truncation (so "typograph*" catches "typography", "typographical" and "typographer"). Accents and capitalisation are ignored when matching. The program compares these terms against the word cloud already computed for every article, which is a ranked list of the 35 words most characteristic of that article. A match on an article's most prominent word counts for more than a match further down the list: the first-ranked word contributes a full point, the second half a point, the tenth a tenth. A term that also appears in the article's title earns an extra half point.

Because no vocabulary can anticipate every way an author might describe a topic, the program then lets articles borrow evidence from their neighbours. Each article receives a reduced share of the scores of the articles whose word clouds most resemble its own, so a piece whose own wording missed the seed terms can still be recognised through the company it keeps. An article is given a subject only if its combined score passes a fixed threshold, and it receives at most three subjects. Articles that meet no threshold are deliberately left without a subject rather than being forced into an ill-fitting one. The whole process is purely arithmetical and uses no external services or AI tools, so the same vocabulary and the same articles will always produce the same result.

The program is designed to be refined over successive runs rather than set up once and left alone. Each run can be made as a trial that changes nothing in the catalog and produces a report. For every subject, the report shows how many articles it attracted, how many of those came through neighbouring articles rather than their own vocabulary, and a few of the strongest examples. It also lists words that are unusually common among a subject's articles compared with the collection as a whole. These are offered as candidate seed terms, which an editor may choose to add to the vocabulary. In time, we will both refine the initial controlled vocabulary and re-run the subject generator to more closely fit the subjects to the articles.

The report closes with a sample of articles that received no subject, each shown with its most prominent words, which often reveals a topic the vocabulary does not yet cover. This allows us to investigate why any single article, identified by its DOI, received or missed a given subject, and the program will set out the words and scores behind the decision. When the results are satisfactory, the assignments are written to the catalog database, and the wiki's subject pages and wiki categories are generated from them. Each assignment is stored with its score, the terms that matched, and whether it arose from the article's own words or from its neighbours, so every subject in the wiki can be traced back to its evidence.

Difficulties with OCR

Several types of errors have been introduced by the various attempts at recognising the text using OCR:

  • inadequate naming: many reviews were entitled "Review" or "Reviews"
  • name concatenation: Alfred Pollard appearing as POLLARDALFRED
  • name separation: "Ed. By Ad Stijnman and Elizabeth Savage" appearing as "Ed. By A d S tijnman and E lizabeth S avage" - caused by misinterpreting small caps

For the first issue, an early solution was to look up each such title in Crossref and use their version of the title. This introduced the remaining issues, which had to be sorted with several runs of a title clean-up program specifically designed to fix those problems.

Another issue arises in the case of Reviews in The Library (all 2919 of them). Unlike regular articles, reviews do not necessarily start at the top of a page or end at the foot of the page. As a consequence, the previous and next reviews had text mixed in with the text of the particular review. This also affects, to a lesser extent, Bibliographical Notes, Society Business, Recent Publications, Correspondence, and even a few Articles. This necessitated building an additional step to (1) pull the text from the PDF, (2) identify where the article of interest starts and ends, and (3) isolate that text and save it in the database. The programs that use the text of each article, such as the program that builds the wordclouds, can then run on a more focussed text sample.

Featured Article of the Day

The Featured Article of the Day on the home page is implemented in Javascript in the user's browser so it uses local midnight to turn over to the next day's selection. A pre-calculated list of articles for every day in the near future is loaded by the browser, today's date is located in the list, and the title, author, publications date, and word cloud for that article is included at the bottom of the home page. Example featured article on 27 September, 2026