From Codd's Relational Model to SQL: The System R Research Lineage
Trace how Codd's relational model, SQUARE and SEQUEL, and IBM's System R turned data independence into a practical query system.
SQL is often described as though it appeared fully formed when someone invented a convenient way to ask a database questions. The primary literature shows a longer and more precise story. Edgar F. Codd proposed a relational model to separate what users meant from how a machine stored data. IBM researchers then explored languages for expressing operations over relations, and the System R project tested whether that model could support a complete, concurrent database with practical performance. The work connected a mathematical model to an engineered system, but those were distinct contributions made over more than a decade.
That distinction matters because the relational model is not SQL, SQL is not the relational model, and System R was explicitly a research project rather than a product. Reading the original papers in sequence reveals how each answered a different problem: representation, query expression, and system implementation.
The problem was dependence on physical organization
Database applications of the period commonly depended on the organization chosen by a particular file or database system. A program could have to know how records were grouped, how links were traversed, or which access path reached a desired item. When storage layouts changed to improve performance or accommodate new data, application code could become entangled with those changes. Codd’s 1970 paper begins from the goal of protecting users and application programs from having to know the internal representation of a data bank.
Codd proposed representing information with n-ary relations. A relation can be understood as a set of tuples described by named attributes, with the model providing rules for describing valid values and relationships. This was not merely a proposal to draw rows and columns on a screen. It was a formal model intended to let a user reason about information independently of its physical arrangement. The paper also discussed normal forms and how redundancy can create consistency problems in a user’s model.
Data independence was the core architectural promise: if a database manager could change its internal representation without forcing users to rewrite their requests, logical work could be insulated from storage engineering. It did not mean that physical design ceased to matter. Storage, indexes, and access methods still affected performance. The proposal was to move that complexity behind an interface rather than make every application encode it.
A model still needs a language
A formal data model does not by itself tell a system how users will retrieve or update information. In a 1971 paper, Codd compared procedural, algebraic, and relational-calculus-based approaches to database sublanguages. He argued that a higher-level relational language could improve intersystem compatibility and standardization because users could describe a desired result rather than dictate a path through stored records.
The distinction between a declarative request and an access procedure is still visible in a query such as:
SELECT department, COUNT(*)
FROM employee
WHERE active = TRUE
GROUP BY department;
This modern illustrative example states a result to compute, not a sequence of disk operations. It is not presented as a literal transcription of an early System R query. The database must still choose how to find qualifying records and aggregate them, but that plan can be selected by the system rather than embedded in every client program.
IBM researchers explored different ways to make that separation usable. SQUARE, described by Raymond Boyce, Donald Chamberlin, Morton Hammer, and W. Frank King in 1973, was a set-oriented data sublanguage for query and update operations. The following year, Chamberlin and Boyce presented SEQUEL, short for Structured English Query Language, as a language for accessing an integrated relational database. Their paper described keyword templates over tabular structures and explained that more complex expressions could be composed from simpler forms. Its stated audience included professional programmers and less frequent database users.
These documents are related stages, not interchangeable names for one frozen language. Codd provided a data model and advocated a class of high-level sublanguages. SQUARE explored one design. SEQUEL described another language for a specific research effort. SQL’s subsequent development inherited from this work, but the 1974 paper should be cited for what it actually describes rather than treated as a complete specification of every later SQL dialect.
System R tested the idea as a database, not a demo query
The System R team had to determine whether relational access could coexist with the features users expected from a database manager. Its 1976 architecture paper describes relational views over shared underlying data and data control facilities including authorization, integrity assertions, triggered transactions, logging and recovery, and support for consistency under shared updates. The paper is explicit that System R was a vehicle for research in database architecture and was not planned as a product at that point.
That scope is historically important. A single-user prototype that can answer a query proves little about multiuser operation, recovery after faults, authorization, or the consequences of concurrent updates. System R addressed these as parts of one design. The logical interface could expose relations and views while a physical layer chose storage structures. Integrity assertions described conditions to preserve. Logging and recovery supplied machinery for recovering a consistent state after failure. Authorization mattered because shared data needed controlled access, not only a way to print a report.
System R also forced the team to confront the cost of hiding physical access details. A declarative query gives the user freedom from specifying a path, but the system must find a good one. The 1979 access-path-selection paper describes how System R selected paths for simple and joined queries expressed in a nonprocedural language. This connects the relational interface to query optimization: the optimizer translates a result-oriented request into a plan using available access paths. A user-facing abstraction only succeeds at scale if the machinery beneath it can choose plans with acceptable performance.
The optimizer therefore was not an optional embellishment on top of SQL. It was one of the mechanisms that made the promise of data independence credible. Applications could ask what rows or values they needed while the database evaluated alternative ways to obtain them. This did not guarantee that every query would be fast, and it did not eliminate the need for indexes or good schema design. It changed who was responsible for choosing the low-level path and made that decision part of the database system.
Evaluation changed the project from architecture to evidence
The 1981 System R history paper describes three principal project phases and discusses what the team learned. Its account separates early work on the SQL interface and a limited single-user implementation from the later multiuser system and evaluation in actual use. The evaluation phase included experience at the IBM San Jose Research Laboratory and other user sites. This is stronger evidence than simply saying that the architecture paper proposed a feature: it documents a research program that built, exercised, and reconsidered a working system.
That distinction also helps with claims about commercial products. The IBM papers establish the technical lineage between the relational model, the SQL research language, and System R’s architecture and evaluation. They do not justify reducing the history of every commercial relational database to a direct copy of one prototype. Research results, staff expertise, subsequent product engineering, standards work, and competing relational projects all formed part of the larger transition.
What changed, and what did not
The lasting shift was not that database storage became simple. It was that applications could express operations against a logical data model while the database manager handled storage access, concurrency, authorization, and recovery. Codd’s abstraction, Chamberlin and Boyce’s language research, and the System R team’s implementation work fit together, but their roles should remain distinct when assigning credit.
Nor did the relational model make all data naturally tabular or make physical performance irrelevant. The original papers are about a formal representation and the engineering needed to make it useful. Their history is best understood as an argument carried by artifacts: Codd’s 1970 model, his 1971 sublanguage analysis, SQUARE and SEQUEL language papers, System R’s design, an access-path paper, and the later project evaluation. Together they show why modern relational databases can expose a stable logical interface without requiring each application to navigate the hardware-level organization itself.
Related:
- How to Trace a Technology’s Lineage Through Patents and Standards Documents
- How to Research a Tech History Topic Using Primary Sources
Sources:
- E. F. Codd, A Relational Model of Data for Large Shared Data Banks (1970), IBM Research publication record
- E. F. Codd, A Data Base Sublanguage Founded on the Relational Calculus (1971), IBM Research
- Boyce, Chamberlin, Hammer, and King, Specifying Queries as Relational Expressions (1973), IBM Research
- Chamberlin and Boyce, SEQUEL: A Structured English Query Language (1974), IBM Research
- Astrahan et al., System R: Relational Approach to Database Management (1976), IBM Research
- Selinger et al., Access Path Selection in a Relational Database Management System (1979), IBM Research
- Chamberlin et al., A History and Evaluation of System R (1981), IBM Research