| Type: | Package |
| Title: | High-Performance Streaming Reader for 'DATASUS' DBC Files |
| Version: | 0.2.0 |
| Date: | 2026-09-30 |
| Description: | A modern, memory-efficient R package for reading and converting DBC files produced by the Brazilian Ministry of Health's 'DATASUS' system. DBC files are DBF databases compressed with the 'PKWare' implode algorithm. Unlike 'read.dbc', this package processes files in streaming mode, never loading the full dataset into RAM. It can export directly to CSV or DBF and provides R helpers for in-memory reading and Parquet conversion. |
| License: | AGPL-3 |
| URL: | https://github.com/GPimentel14/dbcturbo |
| BugReports: | https://github.com/GPimentel14/dbcturbo/issues |
| Encoding: | UTF-8 |
| Depends: | R (≥ 4.0.0) |
| NeedsCompilation: | yes |
| SystemRequirements: | C99 compiler |
| Suggests: | data.table (≥ 1.14.0), arrow (≥ 12.0.0), testthat (≥ 3.0.0), knitr, rmarkdown |
| VignetteBuilder: | knitr |
| Config/roxygen2/version: | 8.1.0 |
| Packaged: | 2026-09-30 18:03:41 UTC; gumercindo |
| Author: | Gumercindo Pimentel Peralta
|
| Maintainer: | Gumercindo Pimentel Peralta <gumercindopimentel@gmail.com> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-10 11:20:20 UTC |
dbcturbo: High-Performance Streaming Reader for 'DATASUS' DBC Files
Description
A modern, memory-efficient R package for reading and converting DBC files produced by the Brazilian Ministry of Health's 'DATASUS' system. DBC files are DBF databases compressed with the 'PKWare' implode algorithm. Unlike 'read.dbc', this package processes files in streaming mode, never loading the full dataset into RAM. It can export directly to CSV or DBF and provides R helpers for in-memory reading and Parquet conversion.
Author(s)
Maintainer: Gumercindo Pimentel Peralta gumercindopimentel@gmail.com (ORCID)
Authors:
Gumercindo Pimentel Peralta gumercindopimentel@gmail.com (ORCID)
Other contributors:
Juliana da Silva (ORCID) [thesis advisor]
Mark Adler (Author of blast.c (PKWare implode decompressor)) [contributor]
Daniela Petruzalek (Author of read.dbc, dbc2dbf.c) [contributor]
Pablo Fonseca (Author of blast-dbf) [contributor]
See Also
Useful links:
Report bugs at https://github.com/GPimentel14/dbcturbo/issues
Decompress a DATASUS DBC file to DBF format
Description
Calls the C streaming engine to decompress a .dbc file into a
standard .dbf file without loading the full payload into RAM.
The compressed payload is processed in 64 KB chunks using the PKWare
Implode algorithm (blast.c, Mark Adler).
Usage
dbc2dbf(input_file, output_file)
Arguments
input_file |
Character string. Path to the source |
output_file |
Character string. Path for the output |
Value
TRUE invisibly on success. On failure, stops with a
descriptive error message.
See Also
Examples
dbc <- system.file("extdata", "sids.dbc", package = "dbcturbo")
dbf <- tempfile(fileext = ".dbf")
dbc2dbf(dbc, dbf)
file.exists(dbf)
unlink(dbf)
Convert all DBC files in a directory to CSV files
Description
Finds DBC files in a directory and converts each one with
dbc_to_csv. When recursive = TRUE, the relative
directory layout beneath input_dir is preserved in
output_dir. Existing output files are never replaced unless
overwrite = TRUE.
Usage
dbc_batch_to_csv(
input_dir,
output_dir,
pattern = "\\.dbc$",
recursive = FALSE,
batch_size = 4096L,
encoding = NULL,
cols = NULL,
overwrite = FALSE,
verbose = FALSE,
workers = 1L
)
Arguments
input_dir |
Character string. Directory containing DBC files. |
output_dir |
Character string. Destination directory for CSV files. Created when it does not exist. |
pattern |
Regular expression used to select input files. Defaults to
|
recursive |
Logical. Search subdirectories? Default |
batch_size |
Integer. Number of records per CSV write batch. Passed to
|
encoding |
Character string or |
cols |
Character vector or |
overwrite |
Logical. Replace existing CSV files? Default |
verbose |
Logical. Print per-file conversion progress? Default
|
workers |
Positive integer. Number of files to convert concurrently on
macOS and Linux. Defaults to |
Value
A data frame with one row per converted file and columns
input_file and output_file. For an empty input directory,
returns a zero-row data frame with those columns.
Examples
input_dir <- tempfile("dbc-input-")
output_dir <- tempfile("dbc-output-")
dir.create(input_dir)
file.copy(system.file("extdata", "sids.dbc", package = "dbcturbo"),
file.path(input_dir, "sids.dbc"))
result <- dbc_batch_to_csv(input_dir, output_dir)
result
unlink(c(input_dir, output_dir), recursive = TRUE)
Inspect the metadata of a DATASUS DBC file without decompressing data
Description
Reads only the DBF header embedded in the .dbc file and returns
field metadata and the total record count. The compressed data payload
is never touched, making this function very fast even for large files.
Usage
dbc_inspect(input_file)
Arguments
input_file |
Character string. Path to the |
Value
A named list with:
fieldsA
data.framewith columnsname(character),type(one-character string:C=text,N=numeric,D=date,L=logical),width(integer),decimals(integer).nrecordsInteger. Total number of records.
See Also
Examples
dbc <- system.file("extdata", "sids.dbc", package = "dbcturbo")
meta <- dbc_inspect(dbc)
meta$nrecords
head(meta$fields)
Convert a DATASUS DBC file directly to CSV (streaming, low-RAM)
Description
Decompresses a .dbc file and writes a UTF-8 CSV to disk.
Records are processed in batches, so peak RAM usage is
O(batch_size * record_width) regardless of file size.
Usage
dbc_to_csv(
input_file,
output_file,
batch_size = 4096L,
encoding = NULL,
cols = NULL,
verbose = FALSE,
progress = NULL,
...
)
Arguments
input_file |
Character string. Path to the source |
output_file |
Character string. Path for the output |
batch_size |
Integer |
encoding |
Character string or |
cols |
Character vector or |
verbose |
Logical. If |
progress |
Function or |
... |
Optional internal arguments. |
Details
The output CSV is prefixed with a UTF-8 BOM so that Microsoft Excel opens it with correct encoding without additional configuration.
Value
TRUE invisibly on success. On failure, stops with a
descriptive error message.
See Also
Examples
dbc <- system.file("extdata", "sids.dbc", package = "dbcturbo")
csv <- tempfile(fileext = ".csv")
dbc_to_csv(dbc, csv)
head(utils::read.csv(csv, fileEncoding = "UTF-8"))
unlink(csv)
Convert a DATASUS DBC file to Parquet format
Description
A high-level convenience wrapper that first writes the .dbc file to
a temporary CSV using the C engine, then converts that CSV into a
.parquet file using the arrow package. It therefore requires
temporary disk space for the CSV and is not an end-to-end streaming writer.
Usage
dbc_to_parquet(
input_file,
output_file,
batch_size = 8192L,
encoding = "CP850",
verbose = FALSE,
progress = NULL
)
Arguments
input_file |
Character string. Path to the source |
output_file |
Character string. Path for the output |
batch_size |
Integer. Passed to |
encoding |
Character string. Source encoding of character fields.
Default |
verbose |
Logical. If |
progress |
Function or |
Details
This is the recommended workflow for Big Data and epidemiological research, as Parquet files are heavily compressed, columnar, and preserve types.
Value
TRUE invisibly on success. Stops if the arrow package
is not installed.
Examples
if (requireNamespace("arrow", quietly = TRUE)) {
dbc <- system.file("extdata", "sids.dbc", package = "dbcturbo")
parquet <- tempfile(fileext = ".parquet")
dbc_to_parquet(dbc, parquet)
arrow::read_parquet(parquet)
unlink(parquet)
}
Read a DATASUS DBC file into an R data frame
Description
High-level convenience wrapper. The native engine decodes small files
directly into R columns without temporary files. Larger files use a
temporary CSV and data.table::fread() (if available) or
utils::read.csv().
Usage
read_dbc(
file,
batch_size = 4096L,
encoding = "CP850",
verbose = FALSE,
cols = NULL,
coerce_types = TRUE,
engine = c("auto", "native", "data.table", "base"),
native_threshold = 50 * 1024^2,
...
)
Arguments
file |
Character string. Path to the |
batch_size |
Integer. Passed to |
encoding |
Character string. Source encoding of character fields.
Default |
verbose |
Logical. Passed to |
cols |
Character vector or |
coerce_types |
Logical. If
|
engine |
Character string selecting the reader. |
native_threshold |
Non-negative number of bytes. In |
... |
Additional arguments forwarded to the CSV reader. |
Details
For files with more than one million records, use dbc_to_csv
directly and load the resulting CSV with arrow::read_csv_arrow() or
data.table::fread() for maximum performance.
Value
A data.frame or data.table (if data.table
is installed).
Examples
dbc <- system.file("extdata", "sids.dbc", package = "dbcturbo")
df <- read_dbc(dbc)
head(df)
# Read a subset of fields.
fields <- dbc_inspect(dbc)$fields$name[1:3]
selected <- read_dbc(dbc, cols = fields)
names(selected)