---
title: "roost(): Time-Unit Aggregation for Surveillance Data"
subtitle: "From daily records to epi curves, with zero-filling and seasonal awareness"
author: "Dr Nicolas Smoll, SCPHU, Sunshine Coast Hospital and Health Service"
date: "`r Sys.Date()`"
output:
  html_document:
    toc: true
    toc_depth: 3
    toc_float: true
    theme: flatly
  pdf_document:
    toc: true
    toc_depth: 3
    number_sections: true
    latex_engine: xelatex
vignette: >
  %\VignetteIndexEntry{roost(): Time-Unit Aggregation for Surveillance Data}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include=FALSE}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>", warning = FALSE, message = FALSE)
library(mudnester)
```

## What `roost()` does

A roost is where many individual birds settle together at dusk — many separate movements resolving into one countable, structured gathering. `roost()` does the same thing to surveillance records: many individual rows resolve into counts at the time unit that matters for the analysis.

The key design feature is **zero-filling**: `roost()` builds a complete calendar grid from the first to the last date in the data, joins real counts onto it, and fills any missing periods with `0`. Epi curves produced from a `roost_tbl` therefore never silently skip empty weeks, which is a common source of misleading visualisations.

---

## Synthetic data

```{r data}
set.seed(42)
n <- 120
df <- data.frame(
  onset_date = as.Date("2024-01-01") + sample(0:364, n, replace = TRUE),
  age        = sample(0:90, n, replace = TRUE),
  pathogen   = sample(c("COVID-19","Influenza A","RSV"), n, TRUE,
                       prob = c(0.45, 0.35, 0.20)),
  icu_flag   = sample(0:1, n, TRUE, prob = c(0.9, 0.1)),
  stringsAsFactors = FALSE
)
df <- preening(df, age_col = "age", scheme = "flucan_sentinel")
```

---

## Available time units

```{r time-units-table, echo=FALSE}
knitr::kable(
  data.frame(
    `time_unit` = c("day","isoweek","fortnight","month","biannual",
                     "quarter","year","epiweek","season","season_year"),
    `Output column type` = c("Date","Date (Monday of week)","Date (first day of fortnight)",
                              "Date (1st of month)","Date (Jan 1 or Jul 1)",
                              "Date (1st of quarter)","Date (Jan 1)","Integer (1–53)",
                              "Character","Character"),
    `Notes` = c("One row per calendar day","ISO 8601 week","14-day intervals from first date",
                 "","H1 = Jan–Jun, H2 = Jul–Dec",
                 "","","Also produces epiyear column",
                 "Hemisphere-aware","Hemisphere-aware; e.g. 'Winter 2024'")
  ),
  col.names = c("time_unit", "Output type", "Notes")
)
```

---

## Examples

### Monthly counts by pathogen

```{r monthly}
monthly <- roost(
  df,
  date_col   = "onset_date",
  time_unit  = "month",
  group_cols = "pathogen"
)
monthly
```

### Epidemiological weeks

`epiweek` also produces an `epiyear` column, so cross-year datasets remain unambiguous.

```{r epiweek}
epi <- roost(df, date_col = "onset_date", time_unit = "epiweek")
head(epi, 6)
```

### Seasons (southern hemisphere)

```{r season}
seasonal <- roost(df, date_col = "onset_date", time_unit = "season_year")
seasonal
```

### Biannual — half-year aggregation

Useful for six-monthly program reporting.

```{r biannual}
bi <- roost(df, date_col = "onset_date", time_unit = "biannual")
bi
```

### Event columns — counting outcomes

Supply `event_cols` to sum binary (0/1) outcome columns alongside the row count.

```{r events}
hosp_counts <- roost(
  df,
  date_col   = "onset_date",
  time_unit  = "month",
  event_cols = "icu_flag",
  group_cols = "pathogen"
)
head(hosp_counts)
```

---

## Stratified aggregation after `preening()`

`preening()` and `roost()` are designed to compose naturally. Age-group columns produced by `preening()` feed directly into `group_cols`.

```{r preening-roost}
age_monthly <- roost(
  df,
  date_col   = "onset_date",
  time_unit  = "month",
  group_cols = c("age_group", "pathogen")
)
head(age_monthly)
```

---

## The `roost_tbl` object

`roost()` returns a `roost_tbl` — a classed tibble. The print method displays metadata automatically.

```{r roost-tbl}
monthly_simple <- roost(df, date_col = "onset_date", time_unit = "month")
monthly_simple   # print.roost_tbl shows the roost_meta footer
```

Metadata survives subsetting:

```{r roost-meta}
sub <- monthly_simple[monthly_simple$n > 5, ]
attr(sub, "roost_meta")$time_unit
```

---

## Zero-filling matters

Without zero-filling, a plot that skips empty weeks can make a declining outbreak look flat or a seasonal upturn look sudden. `roost()` always zero-fills, so you always see the true shape of the curve.

```{r zero-fill-demo}
# Even for a sparse dataset with genuine zero-count periods, every period appears
sparse <- data.frame(onset_date = as.Date(c("2024-01-15","2024-04-20","2024-11-01")))
roost(sparse, date_col = "onset_date", time_unit = "month")
```

---

## What comes next

The `roost_tbl` from `roost()` is the primary input to `bowerbird::roost_plot()` for epi curve visualisation. Before sharing the underlying linelist, consider `molting()` (see `vignette("molting")`).
