| Title: | Secure and Intuitive Access to 'BigDataPE' 'API' Datasets |
| Version: | 0.3.0 |
| Date: | 2026-09-29 |
| Description: | Designed to simplify the process of retrieving datasets from the 'Big Data PE' platform using secure token-based authentication. It provides functions for securely storing, retrieving, and managing tokens associated with specific datasets, as well as fetching and processing data. The data-retrieval engine is provided by the generic 'apifetch' package, which 'BigDataPE' configures for the Big Data PE service. |
| License: | MIT + file LICENSE |
| Encoding: | UTF-8 |
| Depends: | R (≥ 4.1.0) |
| Imports: | apifetch (≥ 0.2.0) |
| URL: | https://strategicprojects.github.io/BigDataPE/, https://github.com/StrategicProjects/BigDataPE |
| BugReports: | https://github.com/StrategicProjects/BigDataPE/issues |
| Suggests: | knitr, rmarkdown, tidyr, dplyr, tibble, httr2 (≥ 1.0.0), testthat (≥ 3.0.0), withr |
| Config/testthat/edition: | 3 |
| VignetteBuilder: | knitr |
| Config/Needs/website: | tidyverse/tidytemplate |
| Config/roxygen2/version: | 8.0.0 |
| NeedsCompilation: | no |
| Packaged: | 2026-10-01 10:50:44 UTC; leite |
| Author: | André Leite |
| Maintainer: | André Leite <leite@castlab.org> |
| Repository: | CRAN |
| Date/Publication: | 2026-10-01 11:50:08 UTC |
Fetch data from the BigDataPE API in chunks
Description
This function retrieves data from the BigDataPE API iteratively in chunks.
It calls the API repeatedly with an advancing offset, stopping when a chunk
comes back empty or total_limit is reached, and combines the chunks into a
single tibble (dropping the API's Mensagem status column). If the API
returns more rows than requested, the result is truncated to total_limit.
Usage
bdpe_fetch_chunks(
base_name,
total_limit = Inf,
chunk_size = 50000L,
query = list(),
verbosity = 0L,
endpoint = "https://www.bigdata.pe.gov.br/api/buscar"
)
Arguments
base_name |
A string specifying the name of the dataset associated with the token. |
total_limit |
An integer specifying the maximum number of records to fetch. Default is Inf (all available data). |
chunk_size |
An integer specifying the number of records to fetch per chunk. Default is 50000;
|
query |
A named list of additional query parameters to filter the API results. Default is an empty list. |
verbosity |
An integer specifying the verbosity level. Values are:
|
endpoint |
A string specifying the API endpoint URL. Default is "https://www.bigdata.pe.gov.br/api/buscar". |
Value
A tibble containing all the data retrieved from the API.
Examples
## Not run:
# Store a token for the dataset
bdpe_store_token("dengue_dataset", "token")
# Fetch up to 500 records in chunks of 100
data <- bdpe_fetch_chunks("dengue_dataset", total_limit = 500, chunk_size = 100)
# Fetch all available data in chunks of 200
data <- bdpe_fetch_chunks("dengue_dataset", chunk_size = 200)
## End(Not run)
Fetch data from the BigDataPE API
Description
This function retrieves data from the BigDataPE API using securely stored tokens associated with datasets.
Users can specify pagination parameters (limit and offset) and additional query filters to customize the data retrieval.
Usage
bdpe_fetch_data(
base_name,
limit = Inf,
offset = 0L,
query = list(),
verbosity = 0L,
endpoint = "https://www.bigdata.pe.gov.br/api/buscar"
)
Arguments
base_name |
A string specifying the name of the dataset associated with the token. |
limit |
An integer specifying the maximum number of records to retrieve per request. Default is Inf (all records).
If set to a non-positive value or |
offset |
An integer specifying the starting record for the query. Default is 0.
If set to a non-positive value or |
query |
A named list of additional query parameters to filter the API results. Default is an empty list. |
verbosity |
An integer specifying the verbosity level. Values are:
|
endpoint |
A string specifying the API endpoint URL. Default is "https://www.bigdata.pe.gov.br/api/buscar". |
Value
A tibble containing the data returned by the API, without the
API's Mensagem status column.
Examples
## Not run:
# Store a token for the dataset
bdpe_store_token("dengue_dataset", "token")
# Fetch 50 records from the beginning
data <- bdpe_fetch_data("dengue_dataset", limit = 50)
# Fetch records with additional query parameters
data <- bdpe_fetch_data("dengue_dataset", query = list(field = "value"))
# Fetch all data without limits
data <- bdpe_fetch_data("dengue_dataset", limit = Inf)
## End(Not run)
Retrieve the token associated with a specific dataset
Description
This function retrieves the authentication token stored
in an environment variable for a specific dataset. If the token is not found,
it returns NULL and prints a warning message instead of throwing an error.
Usage
bdpe_get_token(base_name)
Arguments
base_name |
The name of the dataset (character). |
Value
A string containing the authentication token, or NULL if the token is not found.
Examples
token <- bdpe_get_token("education_dataset")
List all datasets that have stored tokens in environment variables
Description
This function returns a character vector of dataset names
that have their tokens stored in environment variables.
Specifically, it looks for variables that begin with the prefix
"BigDataPE_".
Usage
bdpe_list_tokens()
Value
A character vector of dataset names with stored tokens. If no tokens are found, an empty vector is returned and a message is printed.
Examples
bdpe_list_tokens()
Remove the token associated with a specific dataset
Description
This function removes the authentication token stored in an environment variable for a specific dataset. If the token is not found, it prints a message and does not throw an error.
Usage
bdpe_remove_token(base_name)
Arguments
base_name |
The name of the dataset (character). |
Value
No return value. If the token is found, it is removed. If not, a message is displayed.
Examples
bdpe_remove_token("education_dataset")
Store a token in an environment variable for a specific dataset
Description
This function stores an authentication token for a specific dataset in a system environment variable. The environment variable name is constructed by converting the dataset name to ASCII (removing accents), replacing spaces with underscores, and prefixing it with "BigDataPE_".
Usage
bdpe_store_token(base_name, token, overwrite = FALSE)
Arguments
base_name |
The name of the dataset (character). |
token |
The authentication token for the dataset (character). |
overwrite |
Replace an existing token for this dataset? Default |
Details
If a variable with that name already exists (and is non-empty), the function
will not overwrite it unless overwrite = TRUE (e.g. to replace an expired
token).
Value
No return value, called for side effects.
Examples
bdpe_store_token("education_dataset", "your-token-here")
# Replace an expired token
bdpe_store_token("education_dataset", "new-token", overwrite = TRUE)
bdpe_remove_token("education_dataset")
Constructs a URL with query parameters
Description
This function appends a list of query parameters to a base URL. It is a thin
re-export of apifetch::parse_queries(). Names and values are URL-encoded;
parameters whose value is NULL, NA or the empty string are dropped; a
parameter with several values is repeated (a=1&a=2); and if url already
has a query string, the new parameters are appended to it.
Usage
parse_queries(url, query_list)
Arguments
url |
The base URL to which query parameters will be added. |
query_list |
A named list of query parameters to be added to the URL. |
Value
The complete URL with the query parameters appended.
Examples
parse_queries("https://www.example.com", list(param1 = "value1", param2 = "value2"))