What roost() does

A roost is where many individual birds settle together at dusk — many separate movements resolving into one countable, structured gathering. roost() does the same thing to surveillance records: many individual rows resolve into counts at the time unit that matters for the analysis.

The key design feature is zero-filling: roost() builds a complete calendar grid from the first to the last date in the data, joins real counts onto it, and fills any missing periods with 0. Epi curves produced from a roost_tbl therefore never silently skip empty weeks, which is a common source of misleading visualisations.


Synthetic data

set.seed(42)
n <- 120
df <- data.frame(
  onset_date = as.Date("2024-01-01") + sample(0:364, n, replace = TRUE),
  age        = sample(0:90, n, replace = TRUE),
  pathogen   = sample(c("COVID-19","Influenza A","RSV"), n, TRUE,
                       prob = c(0.45, 0.35, 0.20)),
  icu_flag   = sample(0:1, n, TRUE, prob = c(0.9, 0.1)),
  stringsAsFactors = FALSE
)
df <- preening(df, age_col = "age", scheme = "flucan_sentinel")

Available time units

time_unit Output type Notes
day Date One row per calendar day
isoweek Date (Monday of week) ISO 8601 week
fortnight Date (first day of fortnight) 14-day intervals from first date
month Date (1st of month)
biannual Date (Jan 1 or Jul 1) H1 = Jan–Jun, H2 = Jul–Dec
quarter Date (1st of quarter)
year Date (Jan 1)
epiweek Integer (1–53) Also produces epiyear column
season Character Hemisphere-aware
season_year Character Hemisphere-aware; e.g. ‘Winter 2024’

Examples

Monthly counts by pathogen

monthly <- roost(
  df,
  date_col   = "onset_date",
  time_unit  = "month",
  group_cols = "pathogen"
)
monthly
#> # A tibble: 36 × 3
#>    pathogen month          n
#>    <chr>    <date>     <int>
#>  1 COVID-19 2024-01-01     7
#>  2 COVID-19 2024-02-01     6
#>  3 COVID-19 2024-03-01     3
#>  4 COVID-19 2024-04-01     7
#>  5 COVID-19 2024-05-01     3
#>  6 COVID-19 2024-06-01     3
#>  7 COVID-19 2024-07-01     2
#>  8 COVID-19 2024-08-01     4
#>  9 COVID-19 2024-09-01     4
#> 10 COVID-19 2024-10-01     7
#> # ℹ 26 more rows
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : month 
#>  date_range : 2024-01-01 to 2024-12-25 
#>  group_cols : pathogen 
#>  hemisphere : southern 
#>  n_rows_in  : 120

Epidemiological weeks

epiweek also produces an epiyear column, so cross-year datasets remain unambiguous.

epi <- roost(df, date_col = "onset_date", time_unit = "epiweek")
head(epi, 6)
#> # A tibble: 6 × 3
#>   epiyear epiweek     n
#>     <dbl>   <dbl> <int>
#> 1    2024       1     5
#> 2    2024       2     0
#> 3    2024       3     3
#> 4    2024       4     3
#> 5    2024       5     2
#> 6    2024       6     2
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : epiweek 
#>  date_range : 2024-01-01 to 2024-12-25 
#>  hemisphere : southern 
#>  n_rows_in  : 120

Seasons (southern hemisphere)

seasonal <- roost(df, date_col = "onset_date", time_unit = "season_year")
seasonal
#> # A tibble: 5 × 2
#>   season_year     n
#>   <chr>       <int>
#> 1 Autumn 2024    30
#> 2 Spring 2024    34
#> 3 Summer 2023    22
#> 4 Summer 2024    13
#> 5 Winter 2024    21
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : season_year 
#>  date_range : 2024-01-01 to 2024-12-25 
#>  hemisphere : southern 
#>  n_rows_in  : 120

Biannual — half-year aggregation

Useful for six-monthly program reporting.

bi <- roost(df, date_col = "onset_date", time_unit = "biannual")
bi
#> # A tibble: 2 × 2
#>   biannual       n
#>   <date>     <int>
#> 1 2024-01-01    58
#> 2 2024-07-01    62
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : biannual 
#>  date_range : 2024-01-01 to 2024-12-25 
#>  hemisphere : southern 
#>  n_rows_in  : 120

Event columns — counting outcomes

Supply event_cols to sum binary (0/1) outcome columns alongside the row count.

hosp_counts <- roost(
  df,
  date_col   = "onset_date",
  time_unit  = "month",
  event_cols = "icu_flag",
  group_cols = "pathogen"
)
head(hosp_counts)
#> # A tibble: 6 × 4
#>   pathogen month          n icu_flag
#>   <chr>    <date>     <int>    <int>
#> 1 COVID-19 2024-01-01     7        1
#> 2 COVID-19 2024-02-01     6        0
#> 3 COVID-19 2024-03-01     3        1
#> 4 COVID-19 2024-04-01     7        0
#> 5 COVID-19 2024-05-01     3        0
#> 6 COVID-19 2024-06-01     3        1
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : month 
#>  date_range : 2024-01-01 to 2024-12-25 
#>  group_cols : pathogen 
#>  event_cols : icu_flag 
#>  hemisphere : southern 
#>  n_rows_in  : 120

Stratified aggregation after preening()

preening() and roost() are designed to compose naturally. Age-group columns produced by preening() feed directly into group_cols.

age_monthly <- roost(
  df,
  date_col   = "onset_date",
  time_unit  = "month",
  group_cols = c("age_group", "pathogen")
)
head(age_monthly)
#> # A tibble: 6 × 4
#>   age_group pathogen month          n
#>   <ord>     <chr>    <date>     <int>
#> 1 0-4       COVID-19 2024-01-01     0
#> 2 0-4       COVID-19 2024-02-01     0
#> 3 0-4       COVID-19 2024-03-01     0
#> 4 0-4       COVID-19 2024-04-01     0
#> 5 0-4       COVID-19 2024-05-01     1
#> 6 0-4       COVID-19 2024-06-01     0
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : month 
#>  date_range : 2024-01-01 to 2024-12-25 
#>  group_cols : age_group, pathogen 
#>  hemisphere : southern 
#>  n_rows_in  : 120

The roost_tbl object

roost() returns a roost_tbl — a classed tibble. The print method displays metadata automatically.

monthly_simple <- roost(df, date_col = "onset_date", time_unit = "month")
monthly_simple   # print.roost_tbl shows the roost_meta footer
#> # A tibble: 12 × 2
#>    month          n
#>    <date>     <int>
#>  1 2024-01-01    11
#>  2 2024-02-01    11
#>  3 2024-03-01     7
#>  4 2024-04-01    12
#>  5 2024-05-01    11
#>  6 2024-06-01     6
#>  7 2024-07-01     7
#>  8 2024-08-01     8
#>  9 2024-09-01    12
#> 10 2024-10-01    13
#> 11 2024-11-01     9
#> 12 2024-12-01    13
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : month 
#>  date_range : 2024-01-01 to 2024-12-25 
#>  hemisphere : southern 
#>  n_rows_in  : 120

Metadata survives subsetting:

sub <- monthly_simple[monthly_simple$n > 5, ]
attr(sub, "roost_meta")$time_unit
#> [1] "month"

Zero-filling matters

Without zero-filling, a plot that skips empty weeks can make a declining outbreak look flat or a seasonal upturn look sudden. roost() always zero-fills, so you always see the true shape of the curve.

# Even for a sparse dataset with genuine zero-count periods, every period appears
sparse <- data.frame(onset_date = as.Date(c("2024-01-15","2024-04-20","2024-11-01")))
roost(sparse, date_col = "onset_date", time_unit = "month")
#> # A tibble: 11 × 2
#>    month          n
#>    <date>     <int>
#>  1 2024-01-01     1
#>  2 2024-02-01     0
#>  3 2024-03-01     0
#>  4 2024-04-01     1
#>  5 2024-05-01     0
#>  6 2024-06-01     0
#>  7 2024-07-01     0
#>  8 2024-08-01     0
#>  9 2024-09-01     0
#> 10 2024-10-01     0
#> 11 2024-11-01     1
#> 
#> -- roost_meta --------------------------------------
#>  time_unit  : month 
#>  date_range : 2024-01-15 to 2024-11-01 
#>  hemisphere : southern 
#>  n_rows_in  : 3

What comes next

The roost_tbl from roost() is the primary input to bowerbird::roost_plot() for epi curve visualisation. Before sharing the underlying linelist, consider molting() (see vignette("molting")).