In his chapters of Lord and Novick's Statistical Theories of Mental Test Scores, Birnbaum lays out the mathematical foundations of latent-trait (item-response) models for tests of dichotomously scored items. He develops the logistic item-response models — the two-parameter (2PL) and three-parameter (3PL) forms — in which the probability that a subject of ability answers an item correctly depends on the item's discrimination, difficulty, and (for the 3PL) guessing parameters. Along the way he introduces the item characteristic curve, the item and test information functions, and maximum-likelihood estimation of ability, giving the machinery of modern item-response theory.
"We consider here tests consisting of items each to be scored 0 or 1 … For convenience, we shall refer to the trait in question simply as 'ability'."
This is where item-response theory became a fully specified statistical model rather than a scaling idea. Birnbaum's logistic 2PL/3PL are still the workhorses of large-scale testing (SAT, GRE, licensing exams), and the two moves that matter most are here: letting items differ in discrimination (so not every item is equally informative, unlike Rasch), and formalizing the information function, which is what makes computerized adaptive testing possible (pick the item most informative at the current ability estimate). On the wiki it anchors the general IRT concept above the Rasch special case and connects to the IRT–factor-analysis equivalence (the normal-ogive twin of these logistic curves). It is also a nice reminder that Birnbaum — better known to statisticians for the likelihood principle — did foundational applied work.