| Title: | Lightweight JSON Parsing and Serialization |
| Version: | 0.1.0 |
| Description: | Converts between JSON text and ordinary R vectors and lists through a small, predictable set of functions, backed by vendored 'yyjson' https://github.com/ibireme/yyjson and requiring no system JSON library. Parsing accepts character, raw and file input and reports failures through structured conditions; serialization writes UTF-8 bytes suitable for use directly as an HTTP request body. The type mapping is deliberately narrow and fully documented, so what goes in and what comes out are both predictable. |
| License: | MIT + file LICENSE |
| URL: | https://github.com/pedrobtz/zujson, https://pedrobtz.github.io/zujson/ |
| BugReports: | https://github.com/pedrobtz/zujson/issues |
| Encoding: | UTF-8 |
| Language: | en-GB |
| Imports: | utils |
| Suggests: | jsonlite, knitr, rmarkdown, testthat (≥ 3.2.0), withr |
| Config/testthat/edition: | 3 |
| Config/roxygen2/version: | 8.0.0 |
| NeedsCompilation: | yes |
| Packaged: | 2026-09-18 13:05:38 UTC; pbtz |
| Author: | Pedro Baltazar [aut, cre, cph], YaoYuan [ctb, cph] (author of the bundled yyjson library) |
| Maintainer: | Pedro Baltazar <pedrobtz@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-02 09:10:01 UTC |
zujson: Lightweight JSON Parsing and Serialization
Description
Converts between JSON text and ordinary R vectors and lists through a small, predictable set of functions, backed by vendored 'yyjson' https://github.com/ibireme/yyjson and requiring no system JSON library. Parsing accepts character, raw and file input and reports failures through structured conditions; serialization writes UTF-8 bytes suitable for use directly as an HTTP request body. The type mapping is deliberately narrow and fully documented, so what goes in and what comes out are both predictable.
Author(s)
Maintainer: Pedro Baltazar pedrobtz@gmail.com [copyright holder]
Authors:
Pedro Baltazar pedrobtz@gmail.com [copyright holder]
Other contributors:
YaoYuan (author of the bundled yyjson library) [contributor, copyright holder]
See Also
Useful links:
Report bugs at https://github.com/pedrobtz/zujson/issues
Parse JSON into R
Description
json_parse() turns JSON text, or the raw bytes of an HTTP response body,
into ordinary R vectors and lists. json_parse_raw() and
json_parse_file() are the same parser reached directly, for when the input
type is already known.
Usage
json_parse(x, simplify = TRUE, data_frame = FALSE)
json_parse_raw(x, simplify = TRUE, data_frame = FALSE)
json_parse_file(path, simplify = TRUE, data_frame = FALSE)
Arguments
x |
For |
simplify |
How to simplify JSON arrays to atomic vectors: one of
|
data_frame |
Whether an array of objects becomes a data frame. |
path |
Path to a file containing JSON. |
Value
The R object the JSON maps to, by the table above: a JSON object
becomes a named list, an array an atomic vector when its elements share
a kind and a list when they do not, a string a length-1 character, a
number a length-1 integer or double, true/false a length-1
logical, and null becomes NULL. The class of the result therefore
depends on the JSON and on both arguments: simplify decides the
mixed-array case, and data_frame = TRUE turns an array of objects into a
data.frame. All three functions return the same value for the same JSON
and differ only in where the bytes come from.
Type mapping
JSON objects always become named lists. JSON arrays become an atomic vector when every element agrees on a type and a list otherwise:
| JSON | R |
{"a": 1} | named list |
[1, 2, 3] | integer |
[1, 2.5] | double |
[true, false] | logical |
["a", "b"] | character |
[1, "a"] | list (no common type without coercing) |
[1, {}] | list (nested container) |
[], [null, null] | logical |
null | NULL
|
Numbers become integer when they fit in R's 32-bit integer and double
otherwise. A number too large for any finite double becomes Inf rather
than an error: RFC 8259 sets no limit on the magnitude of a number, so
1e309 is valid JSON, and the value is the one as.numeric() gives the
same token. A null inside an array being simplified becomes NA; a null
anywhere else becomes NULL. Strings arrive as UTF-8.
Simplification modes
simplify picks what happens to an array whose elements do not share a
kind. The three modes agree everywhere else, including the promotions
within the numeric family ([true, 1] is c(1L, 1L) in all of them):
| mode | [1, "a"] becomes |
"preserve", or TRUE (default) | list(1L, "a") |
"coerce" | c("1", "a") |
"none", or FALSE | list(1L, "a"), and every other array is a list too
|
preserve keeps the type and gives up the uniform shape, on the view that a
field which is usually a number and occasionally a string is a bug worth
seeing. coerce is the opt-in for callers who would rather have the vector:
it follows R's own promotion, so the strings are exactly what
as.character() produces. A nested object or array is never coerced away.
Data frames
data_frame = TRUE turns any non-empty array whose elements are all
objects into a data frame. Columns are the union of the keys in the order
first seen, so the result is rectangular however ragged the records are. A
record missing a key contributes NA in an atomic column and NULL in a
list column – the same as an explicit JSON null in either case, because
"absent" and "null" are one thing once the value is in a column. Each column
is then simplified with the active simplify mode, so a column of mixed
kinds is a list column under preserve and a character column under
coerce. It applies wherever such an array appears, however deeply nested.
simplify = "none" takes precedence: it is the mode that guarantees every
JSON array arrives as an R list, and a data frame is not one. The two
options are not combined, and data_frame = TRUE is ignored under it.
Two limits apply, both raising a structured condition rather than a bare
error. An object with the same key twice cannot become a row – a column has
one cell per record – so it raises zujson_parse_error instead of silently
keeping one of the values; plain parsing, which puts both in a list, is
unaffected. And because the frame is rectangular, its size is set by the
union of the keys rather than by the length of the body: records that share
no keys at all would ask for one cell per record per record, and no limit
on the size of the body can stand in for a limit on that, because the growth
is quadratic in it. More than zujson_info()$max_df_cells cells raises
zujson_limit_error. The default is 50 million – clear of any real tabular
response, and far below what a hostile one reaches – and
options(zujson.max_df_cells = ) changes it for the session.
zujson_info()$max_df_cells reports the limit in force.
Setting it larger than any frame that could be built switches the check off;
the option is clamped to what an R_xlen_t holds, which is what
zujson_info()$max_df_cells then reports. Switched off, nothing bounds the
allocation but the machine, and a body large enough to exhaust memory raises
R's own allocation error rather than a zujson_error.
Nesting deeper than 1000 levels is rejected with a zujson_depth_error,
which is what makes the parser safe to point at an untrusted response body.
A leading UTF-8 byte order mark is ignored rather than rejected: RFC 8259
forbids emitting one but allows ignoring it, and real APIs emit them. The
bare literals Infinity, -Infinity and NaN are not JSON and stay
rejected, which is a separate question from the magnitude of a number that
is written as one.
Two further things are valid JSON and still cannot become R values, and
both raise zujson_parse_error rather than a bare error, so a caller
handling zujson_error catches them alongside everything else: a string or
key containing an escaped NUL (\u0000), which no R string can hold, and
one longer than .Machine$integer.max bytes. json_parse_file()
additionally raises zujson_io_error when the file cannot be read at all,
which is a different problem from its contents not being JSON.
See Also
json_write() for the other direction, json_validate() to check
without building a result.
Examples
json_parse('{"ok": true, "ids": [1, 2, 3]}')
# arrays that have no common type stay lists
json_parse('[1, "a"]')
# raw bytes, as an HTTP response body arrives
json_parse(charToRaw('{"ok": true}'))
# no simplification at all
json_parse('[1, 2, 3]', simplify = FALSE)
# coerce across kinds instead of keeping the type
json_parse('[1, "a"]', simplify = "coerce")
# a number past the range of a double is Inf, not a parse failure
json_parse("[1e309]")
# an array of records, as a data frame
json_parse('[{"id":1,"nm":"a"},{"id":2,"nm":"b"}]', data_frame = TRUE)
json_parse_raw(charToRaw('[1, 2, 3]'))
path <- tempfile(fileext = ".json")
writeLines('{"a": [1, 2]}', path)
json_parse_file(path)
unlink(path)
Parse and write NDJSON
Description
NDJSON (application/x-ndjson, also called JSON Lines) is one JSON value per
line. json_parse_ndjson() turns such a body into a list of records;
json_write_ndjson() and json_write_ndjson_raw() go the other way.
Usage
json_parse_ndjson(x, simplify = TRUE, data_frame = FALSE)
json_write_ndjson(x, auto_unbox = TRUE)
json_write_ndjson_raw(x, auto_unbox = TRUE)
Arguments
x |
For |
simplify |
Passed through to |
data_frame |
Passed through to |
auto_unbox |
Passed through to |
Details
Each line is parsed exactly as json_parse() would parse it, and each record
is written exactly as json_write() would write it, so the type mappings
documented there apply unchanged. The result of parsing is always a list,
one element per record, however uniform the records are — records are
independent documents, and an NDJSON body is not a table.
Blank lines are skipped rather than treated as records, and \r\n line
endings are accepted. A parse failure reports the line number, because
"invalid JSON at byte 41827" is not useful in a body of 10,000 records.
Value
json_parse_ndjson() returns a list with one element per record,
each being what json_parse() returns for that line. json_write_ndjson()
returns a length-1 character vector and json_write_ndjson_raw() a raw
vector of UTF-8 bytes; both hold one record per line and end with a
trailing newline, so appending another record is always valid.
Why line framing is safe
A raw newline byte cannot appear inside a JSON string — it must be escaped as
\\n — and zujson's writer always escapes it. A serialized record therefore
never contains a bare newline, so splitting on newlines can never cut a
record in half. This is the property the format rests on.
It is also why there is no pretty argument: indented JSON contains
newlines, and a newline inside a record is exactly what NDJSON framing cannot
survive. Pretty-printed NDJSON is corrupt, not prettier.
See Also
json_parse() and json_write() for single documents.
Examples
body <- '{"id":1,"ok":true}\n{"id":2,"ok":false}\n'
json_parse_ndjson(body)
# a data frame is one record per row
cat(json_write_ndjson(data.frame(id = 1:2, nm = c("a", "b"))))
# round trip
recs <- list(list(a = 1L), list(b = "x"))
identical(json_parse_ndjson(json_write_ndjson(recs)), recs)
# bytes, ready to be an HTTP request body
json_write_ndjson_raw(list(list(a = 1L), list(a = 2L)))
Check whether input is valid JSON
Description
Parses x and reports whether it succeeded, without building an R result
and without raising a condition. Use it to decide whether a response body is
worth parsing; use json_parse() when a failure should be an error you can
read.
Usage
json_validate(x)
Arguments
x |
A single string of JSON text, or a raw vector of JSON bytes. |
Details
This answers "is this valid JSON", which is very nearly but not exactly
"will json_parse() succeed". The two come apart on JSON that is valid and
still cannot become an R value, of which there are two kinds:
a string or key R cannot hold: one containing an escaped NUL (
\u0000), or one longer than.Machine$integer.maxbytes.json_parse()raiseszujson_parse_error.nesting deeper than 1000 levels. RFC 8259 leaves any depth limit to the implementation, so such a document is still valid JSON.
json_parse()raiseszujson_depth_error, because its recursive tree builder has a C stack to protect; the reader underneath it is iterative and does not.
This function returns TRUE for both, and is right to: it is asked whether
the bytes are JSON, not whether this package can materialise them. Code that
must not fail should handle the condition from json_parse() rather than
pre-screening with this.
Value
A length-1 logical, never NA: TRUE when x is well-formed
JSON by the reader's rules, FALSE otherwise. NA_character_ is FALSE
rather than NA. A TRUE is not a promise that json_parse() will
succeed – the two cases above are valid JSON that this package cannot
turn into an R value.
Examples
json_validate('{"a": 1}')
json_validate('{"a": }')
json_validate(charToRaw("[]"))
Serialize R as JSON
Description
json_write() renders an R object as JSON text. json_write_raw() returns
the same JSON as UTF-8 bytes, which is the form an HTTP request body wants.
Usage
json_write(x, pretty = FALSE, auto_unbox = TRUE)
json_write_raw(x, pretty = FALSE, auto_unbox = TRUE)
Arguments
x |
The R object to serialize. |
pretty |
Whether to indent the output. |
auto_unbox |
Whether a length-1 atomic vector becomes a bare JSON
scalar ( |
Value
json_write() returns a length-1 character vector holding the
JSON text. json_write_raw() returns the same document as a raw vector
of UTF-8 bytes, which is what an HTTP request body wants, so that
direction needs no conversion step.
Type mapping
| R | JSON |
NULL | null |
NA, NaN, Inf | null |
list(a = 1, b = 2) | {"a":1,"b":2} |
list(1, 2) | [1,2] |
c(a = 1, b = 2) | {"a":1,"b":2} |
1:3 | [1,2,3] |
"a" | "a" (see auto_unbox) |
I("a") | ["a"] |
factor | its level, as a string |
Date | "YYYY-MM-DD" |
POSIXct | "YYYY-MM-DDTHH:MM:SSZ", in UTC |
data.frame | array of one object per row |
list() | [] |
structure(list(), names = character()) | {}
|
A vector or list becomes a JSON object when every element is named, and
an array otherwise. Partial names would produce keys like "", which is
valid JSON but almost never intended, so a partially named vector is written
as an array and its names are dropped.
Data frames are written row-oriented, because that is what an HTTP API means
by a table. To get the column-oriented form instead, strip the class first:
json_write(as.list(df)).
Every kind of missing value becomes null: JSON has no NA, and NaN and
Infinity are not JSON either. POSIXct is always written as UTC no matter
what its tzone says, since it is the same instant either way, and
sub-second parts are dropped. Both are doubles in R and so reach much
further than a timestamp can be written: outside 0000-01-01 to
9999-12-31, the four-digit years ISO 8601 allows, they raise a
zujson_write_error rather than print a year no API can parse.
A string whose Encoding() is "bytes" also raises a zujson_write_error.
That marking is R saying it does not know the encoding, and JSON text is
UTF-8 by definition, so there is nothing to convert from. Every other
encoding R tracks is converted on the way out.
Doubles are written compactly and always read back as the same number, so
0.1 is 0.1 rather than 0.10000000000000001. A double that is a whole
number is written without a decimal point at all — R has no integer literal,
so 1 is a double, and 1.0 is rejected by a schema expecting an integer.
Complex vectors, raw vectors, functions, environments and POSIXlt have no
sensible JSON form and raise a zujson_unsupported_type error rather than
being guessed at. (POSIXlt is a list of 11 broken-down time fields; convert
it with as.POSIXct() first.) A vector carrying an unrecognised class is
written as its underlying type, which for a matrix means its values in
column-major order with dim dropped — matrices are not turned into nested
arrays in this version.
See Also
json_parse() for the other direction.
Examples
json_write(list(name = "ada", ids = 1:3, ok = TRUE))
# length-1 vectors unbox by default; I() keeps them arrays
json_write(list(tag = "x", tags = I("x")))
# every atomic vector stays an array
json_write(list(tag = "x"), auto_unbox = FALSE)
cat(json_write(list(a = 1, b = list(c = 2)), pretty = TRUE))
json_write(data.frame(id = 1:2, nm = c("a", "b")))
# bytes, ready to be an HTTP request body
json_write_raw(list(q = "search"))
Report what this build of zujson contains
Description
Returns the package version, the version of the vendored yyjson sources it was compiled against, and the nesting depth limit the parser and serializer both enforce.
Usage
zujson_info()
Value
A named list with zujson, yyjson, max_depth and
max_df_cells. max_depth is compiled in and fixed; max_df_cells is
the limit in force, which options(zujson.max_df_cells = ) changes.
Examples
zujson_info()