JSON vs YAML vs XML: Differences and When to Use Each
Drafted with AI assistance. Every command and code example was run and its output checked before publication. How guides are made
JSON, YAML and XML can all describe the same structured data, but they are built for different jobs. JSON is a small, strict format with six value types and a parser in every language, which makes it the default for APIs and data exchange. YAML is a superset of JSON designed for people to read and edit, so it suits configuration files, at the cost of indentation rules and implicit typing that cause bugs. XML is a markup language: verbose, typeless without a schema, but the only one of the three that handles documents with mixed text and markup, namespaces and mature schema validation.
What are the main differences between JSON, YAML and XML?
| Feature | JSON | YAML | XML |
|---|---|---|---|
| Specification | RFC 8259 | YAML 1.2.2 (many tools still follow 1.1) | XML 1.0 (W3C) |
| Structure | Braces and brackets | Indentation (or JSON-style flow syntax) | Nested start and end tags, plus attributes |
| Built-in types | String, number, boolean, null, object, array | JSON's types; parsers also resolve dates, octal and other forms depending on version | None: all content is text unless a schema assigns types |
| Comments | No | # to end of line | <!-- ... --> |
| Schema language | JSON Schema | None of its own; JSON Schema is commonly used | XSD, also DTD and RELAX NG |
| Mixed text and markup | No | No | Yes |
| References within a file | No | Anchors and aliases | Entities |
| Multiple documents per file | No (use JSON Lines) | Yes, separated by --- | No, one root element |
| Browser support | JSON.parse built in | Needs a library | DOMParser built in |
| Typical use | APIs, data exchange, storage | Configuration, CI pipelines, Kubernetes | Documents, feeds, SOAP, office file formats, enterprise integration |
What does the same record look like in each format?
Here is one staff record in all three. The JSON and YAML versions were parsed with JSON.parse and js-yaml 4 and compared with Node's assert.deepStrictEqual: they produce identical objects. The XML version was parsed with the browser's DOMParser in Chromium and only matched after code converted the text values to a number, a boolean and null, which is the main difference in practice.
JSON
{
"id": 1042,
"name": "Ada Lovelace",
"active": true,
"roles": ["admin", "editor"],
"manager": null
}
YAML
# Staff record
id: 1042
name: Ada Lovelace
active: true
roles:
- admin
- editor
manager: null
XML
<?xml version="1.0" encoding="UTF-8"?>
<!-- Staff record -->
<user>
<id>1042</id>
<name>Ada Lovelace</name>
<active>true</active>
<roles>
<role>admin</role>
<role>editor</role>
</roles>
<manager/>
</user>
Notice what each one needs that the others do not. XML must name the root (user) and each list item (role), because it has no anonymous objects or arrays. JSON cannot hold the comment. YAML needs no quotes here, but only because none of these values looks like another type.
Which format is smallest?
For the record above, measured in bytes without the comment and XML declaration: minified JSON is 89 bytes, the YAML as shown is 83 bytes, indented JSON is 112 bytes and XML with all whitespace removed is still 134 bytes. XML pays for closing tags that repeat every name; YAML saves the quotes and braces but depends on indentation, so it cannot be minified onto one line without switching to its JSON-like flow style.
For data sent over a network the difference matters less than it looks: HTTP compression such as gzip removes much of the repetition, and parsing speed and library support usually matter more than raw size. The minification guide covers this in more detail.
How do the data types differ?
JSON has exactly six kinds of value, and what you see is what you get: "1042" is a string and 1042 is a number. It has no date type, so dates travel as strings, usually in ISO 8601 form, and integers larger than 253 lose precision in JavaScript.
YAML decides the type of an unquoted value by what it looks like, and the rules changed between versions. YAML 1.1 treated yes, no, on and off as booleans and 0644 as octal; YAML 1.2's core schema does not. Parsers still disagree. The same input loaded with js-yaml 4.3.2 and with PyYAML 6.0.3 (which follows YAML 1.1) gives:
| YAML value | js-yaml 4 (JavaScript) | PyYAML 6 safe_load (Python) |
|---|---|---|
no | String "no" | False |
NO (a country code) | String "NO" | False |
on | String "on" | True |
0644 | Number 644 | Number 420 (octal) |
0o644 | Number 420 | String "0o644" |
1.10 | Number 1.1 | Number 1.1 |
12:30 | String "12:30" | Number 750 (base 60) |
2026-10-05 | Date object | datetime.date |
The fix is to quote any value that must remain a string, such as version numbers, postcodes, country codes and file modes. The YAML pitfalls guide covers these rules, indentation and multi-line strings in depth.
XML has no data types at all in the document itself. <active>true</active> contains the text true, and the application decides what it means. An XSD schema can declare that an element is an xs:integer or xs:boolean, and schema-aware tools then validate and convert it. XML also has two places to put data, elements and attributes, which have no equivalent in JSON or YAML, and it has no standard way to write an array or null; repeated elements and an empty or xsi:nil element are conventions.
Which formats support comments?
YAML (# comment) and XML (<!-- comment -->) do; JSON does not, and a strict parser rejects // or /* */. That alone is often why configuration ends up in YAML. Comments are also the first thing lost in conversion: every parser discards them, so YAML or XML turned into JSON and back has lost its comments for good. If you must annotate JSON, use an ordinary field such as "_comment"; the JSON syntax errors guide explains what else strict parsers reject.
How do you validate each format against a schema?
- JSON: JSON Schema describes required properties, types, formats and value ranges, and has validators for most languages. OpenAPI uses it to describe API payloads.
- YAML: has no schema language of its own. Because a YAML document loads to the same data model as JSON, tools validate it with JSON Schema after parsing, which is how editors check files such as CI pipeline definitions.
- XML: XML Schema (XSD) is a W3C standard with typed elements and attributes, and older DTDs are still found. Namespaces let one document combine vocabularies without name clashes, which is why standards bodies and enterprise systems still rely on it.
Well-formed is not the same as valid. All three formatters on this site check that a document parses; none of them validate against a schema.
How does tooling compare?
JSON has the widest support: JSON.parse is built into every browser and JavaScript runtime, Python ships the json module, and command-line tools such as jq are common. XML is also built in almost everywhere (DOMParser in browsers, xml.etree.ElementTree in Python) and comes with XPath for queries and XSLT for transformations. YAML always needs a library: js-yaml in JavaScript, PyYAML or ruamel.yaml in Python. Python's standard library has no YAML parser.
Security differs too. XML parsers that expand external entities are open to XXE attacks, described by OWASP, and the Python XML documentation describes entity-expansion attacks such as "billion laughs". Some YAML loaders can construct arbitrary objects from tags, so use the safe variant (yaml.safe_load in PyYAML; js-yaml 4's load is safe by default). JSON parsers only ever produce plain data.
When should you use JSON, YAML or XML?
APIs and data exchange: JSON
Choose JSON when programs talk to programs: REST and GraphQL APIs, message queues, browser storage, logs (as JSON Lines) and anything a web page fetches. It is unambiguous, fast to parse and native to JavaScript. Pretty-print it when people need to read it; see how to pretty-print JSON.
Configuration edited by people: YAML
Choose YAML when people write the file by hand and comments matter: CI pipelines, Kubernetes manifests, Docker Compose files, application settings. Quote strings that could be read as something else, indent with spaces only, and validate the result. If the configuration is generated by a program, JSON is often the safer choice, and any JSON document is also valid YAML 1.2.
Documents and mixed content: XML
Choose XML when the data is a document, with text that contains markup (<p>Hello <b>world</b></p>), when you need namespaces or XSD validation, or when an existing standard requires it: RSS and Atom feeds, SVG, SOAP services, office document formats and many industry exchange formats. Neither JSON nor YAML can represent interleaved text and elements without inventing a convention.
What is lost when you convert between them?
JSON and YAML convert cleanly in both directions, apart from YAML features JSON lacks. XML is different, because its model (elements, attributes, text, order) does not map one to one onto objects and arrays. These are the results from the site's converters, run on the record above:
- JSON to YAML: nothing is lost. The output for the record matches the YAML above minus the comment. Strings that a YAML 1.1 parser would misread, such as
"no","0644"or"1.10", are written with quotes. - YAML to JSON: comments disappear, anchors and aliases are expanded into copies, and values are typed as js-yaml 4 resolves them, including dates, which become ISO timestamp strings.
- JSON to XML: types are lost, since
1042andtruebecome text. The object gets a<root>element, each array becomes repeated elements named after the key (two<roles>elements), andnullbecomes an empty element. Keys that start with-, such as"-id", are written as attributes. - XML to JSON: every value comes back as a string (
"id": "1042","active": "true"), the empty<manager/>becomes""rather thannull, attributes go under an"@attributes"key, and comments are dropped. A repeated element becomes an array, but the same element appearing once becomes a plain value, so the shape depends on the data. Mixed content loses its order:<p>Hello <b>world</b>, again</p>becomes{"b": "world", "#text": "Hello , again"}.
Because of this, a JSON-XML-JSON round trip does not give back the original document. If you need a lossless mapping, define it explicitly for your data, ideally from a schema, rather than relying on a generic converter. For tabular data, CSV is another option, with its own traps described in CSV pitfalls.
Quick reference
- Machine-to-machine data: JSON. Strict, typed, universally supported, no comments.
- Hand-written configuration: YAML. Comments and clean syntax; quote ambiguous values and use a YAML 1.2 parser where possible.
- Documents, mixed content, namespaces, XSD: XML.
- Schemas: JSON Schema for JSON and YAML; XSD for XML.
- Size: XML is the largest; YAML and minified JSON are close.
- Conversion: JSON and YAML convert cleanly apart from comments; anything to or from XML loses types or structure unless you define the mapping.
- Security: disable external entities in XML parsers; use safe loaders for YAML.
Frequently asked questions
Is JSON valid YAML?
Under YAML 1.2, yes: the specification was revised so that JSON is a subset, and js-yaml 4 parses a JSON document to the same data. YAML 1.1 parsers such as PyYAML accept ordinary JSON too, but the older specification did not guarantee it.
Is YAML slower to parse than JSON?
Generally yes, because the YAML grammar is far larger and parsers must resolve indentation and implicit types, while JSON parsers are small and often built into the runtime. For configuration read once at start-up it rarely matters; for high-volume data exchange, use JSON.
Should a new API use XML or JSON?
JSON, unless your consumers or an industry standard require XML. JSON maps directly onto the objects and arrays in your code, and browsers parse it natively. XML remains the right choice where documents, namespaces or XSD-based contracts are part of the requirement.
Can XML attributes be converted to JSON?
Only by convention, because JSON has no attributes. Converters invent a marker: the site's XML to JSON tool puts them under "@attributes", while its JSON to XML tool turns keys beginning with - into attributes. Other libraries use @id or _attributes, so check which one your code expects.
Why did "no" in my YAML file turn into false?
Your parser follows YAML 1.1, where words such as no, off and on are booleans. Quote the value ("no") or use a YAML 1.2 parser such as js-yaml 4, which reads it as a string.