Skip to content
Tech HistoryDeep Dive Published Updated 7 min readViews unavailable

JSON: From JavaScript Data Notation to an Independent Interchange Standard

Trace JSON from JavaScript object notation through RFC 4627 and ECMA-404, clarifying its syntax, semantics, interoperability, and relationship to JavaScript.

JSON is now one of the most familiar formats for exchanging structured data, but its history is often summarized too casually as “JavaScript objects sent over the network.” JSON was derived from JavaScript’s object-literal notation, yet it developed into a language-independent text format with its own grammar and standards. That distinction is crucial: JSON resembles a subset of JavaScript syntax, but JSON text is not JavaScript source code and should not be evaluated as one.

The format’s success came from a small set of readable structural forms: objects, arrays, strings, numbers, booleans, and null. This compact syntax fit web applications and could be generated or parsed by many programming languages. Standardization later clarified the allowed syntax and interoperable behavior, while leaving semantic meaning to application protocols.

The Web needed a simple data representation

Web applications increasingly needed to exchange structured values between browsers and servers. HTML handled documents and XML offered a mature markup ecosystem, but some applications wanted a lighter syntax that mapped easily to common programming-language values. JSON’s braces, brackets, strings, and literals were familiar to JavaScript developers and small enough to process in many environments.

Douglas Crockford popularized and specified JSON as an interchange format in the early 2000s. The json.org site introduced the grammar and explanatory material; the IETF later standardized the application/json media type and a formal definition in RFC 4627, published in 2006. These stages should be separated: a format’s invention and early documentation, its deployment by developers, and its formal standardization are different historical events.

JSON’s name expands to JavaScript Object Notation, but it is not restricted to the browser language. Ecma International’s ECMA-404 describes it as a lightweight, text-based, language-independent syntax derived from ECMAScript. The standard intentionally defines syntax only. It does not decide what a property named amount means or how a consuming language must map every value into its runtime types.

A deliberately small grammar lowered the barrier to use

JSON values are built from objects, arrays, strings, numbers, true, false, and null. Objects provide name/value associations; arrays preserve an ordered sequence. These forms are enough to represent many nested application structures without introducing a general markup language or a large family of declarations.

The simplicity is relative, not absolute. A parser must still enforce quotation, escapes, numeric syntax, structural delimiters, and valid Unicode handling. JSON does not permit comments, trailing commas, functions, undefined, NaN, or Infinity in the standardized grammar. Some programming environments accept extensions, but a producer that emits them may no longer be producing interoperable JSON.

The format’s design goals described in RFC 8259 include minimality, portability, textual representation, and a relationship to JavaScript. A small grammar made it straightforward to embed JSON parsers in browsers, services, databases, command-line tools, and programming libraries. Ease of parsing did not remove the need to validate schemas or enforce application-level constraints.

RFC 4627 gave JSON a formal Internet identity

RFC 4627, authored by Crockford, described the JSON data interchange format and registered the application/json media type. It made the format legible to Internet protocol designers and gave applications a common label for JSON bodies. The RFC specified a grammar based on a subset of JavaScript and documented encoding and security considerations relevant to early implementations.

The initial RFC had constraints that later specifications changed. RFC 4627 required a top-level object or array. RFC 7159 revised the description in 2014 and allowed any JSON value at the top level. RFC 8259, published in 2017, obsoleted RFC 7159 and emphasized interoperability while aligning its definition with Ecma’s ECMA-404. Current software and historical documents can therefore differ about top-level scalars; an article must identify which specification it is describing.

RFC 8259 also documents a standards coordination problem: the IETF and Ecma International maintain descriptions of JSON, and their texts are intended to agree. The RFC explains the normative relationship and the expectation that the organizations coordinate any changes. This is a small but significant example of a widely deployed format needing more than one standards body to keep its definition aligned.

JSON is not a safe substitute for parsing

Because JSON resembles JavaScript source, a naïve implementation might pass a response to eval. That is dangerous and semantically incorrect. JSON is a data grammar; JavaScript source can execute expressions, invoke functions, and have behavior beyond JSON values. A conforming JSON parser accepts only the data syntax and returns values according to a documented mapping.

The security issue is not hypothetical. Early web programming patterns used JSON-like responses in ways tied to script inclusion and browser execution, leading to cross-site data exposure risks. RFCs discuss implementation security, and modern APIs should set correct content types, validate inputs, avoid unsafe evaluation, and apply appropriate origin and authorization controls. The format itself does not authenticate a sender or guarantee that data is trustworthy.

Parsing is also different from validation. A syntactically correct object may violate an application’s required fields, type constraints, size limits, or business rules. JSON Schema and application code can define those additional constraints, but they are separate layers. A data exchange succeeds only when producers and consumers agree about both the syntax and the meaning of fields.

Relationship to ECMAScript and the JavaScript runtime

JavaScript object literals and JSON overlap, but the standards are not identical. JSON strings require double quotes; object keys must be strings; only the standardized numeric and literal forms are allowed; and values are restricted to JSON’s finite grammar. JavaScript supports additional constructs and runtime values. The JSON.parse and JSON.stringify APIs define mappings between JSON text and JavaScript values, including behavior for values that JSON cannot represent directly.

This relationship helped browser adoption. JavaScript programs could parse and construct JSON without an entirely separate runtime, while server-side languages implemented the same syntax independently. But the mapping can lose information: JavaScript undefined is not a JSON value, and numbers may not represent arbitrary-precision integers exactly. Application protocols must account for such cross-language differences.

The format is intentionally silent about schemas, object identity, comments, and cyclic references. A JSON tree cannot directly express a pointer cycle, and repeated objects become repeated values unless an application specifies conventions. Developers sometimes layer custom encodings on top, but those conventions are not part of JSON and can reduce interoperability with generic parsers.

JSON’s reach grew through APIs and tooling

REST-style web APIs commonly adopted JSON because it was easy to inspect, produce, and consume. Browsers could display its text; developers could compare payloads in logs; and server libraries could map objects to and from the syntax. Mobile applications, configuration files, package metadata, and command-line interfaces also adopted JSON for similar reasons.

The availability of parsers across languages reinforced the network effect. Once a format has stable grammar and ubiquitous tooling, organizations can use it without committing to one vendor’s binary representation. At the same time, JSON is not necessarily the best choice for every workload. Large numerical datasets, streaming records, strict schema evolution, or human-edited configuration may call for different formats or additional validation.

JSON’s popularity also produced divergent expectations. One developer might treat member order as irrelevant; another might depend on it. Duplicate object names are syntactically accepted in some contexts but discouraged because implementations can interpret them differently. RFC 8259 documents interoperability recommendations intended to avoid such ambiguities. Producers should generate conservative JSON and consumers should not assume undocumented behavior.

What standardization did and did not guarantee

ECMA-404 defines the syntax of valid JSON texts and deliberately does not assign semantics. RFC 8259 defines the Internet-facing format and its interoperability recommendations, while the media type identifies JSON representations in protocols. Application specifications still need to state field meanings, security rules, versioning, error behavior, and data constraints.

That layering helps explain JSON’s durability. A small core syntax can be reused in many contexts while protocols provide domain-specific contracts. But it also means that “valid JSON” is not equivalent to “valid request,” “safe input,” or “compatible payload.” Teams must test their schemas, number ranges, encoding expectations, and versioning behavior.

The evidence supports a specific history: JSON emerged from a notation familiar to JavaScript programmers, gained early documentation and practical use through Crockford’s work, was described in RFC 4627 in 2006, and is now specified in parallel by IETF RFC 8259 and Ecma’s ECMA-404. Its achievement was not simply compact syntax; it was a small, inspectable data format whose formal boundary allowed independent implementations to interoperate without running the same programming language.

Related:

Sources:

Comments