Skip to content

Repository files navigation

brfssdata brfssdata hex logo: an ECG pulse trace sweeping across a purple hexagon

CRAN status Downloads R-CMD-check Codecov test coverage Integration Lifecycle: stable License: MIT DOI

brfssdata gives R users direct, reproducible access to annual microdata from the CDC Behavioral Risk Factor Surveillance System (BRFSS), the largest continuously conducted health telephone survey in the world.

Survey years are processed once into compact parquet files and hosted as GitHub release assets. The package downloads each requested year a single time into a local cache, queries it through DuckDB (so pulling a handful of variables from a 300-plus column survey is fast), and hands back either a tibble or a ready-made srvyr survey-design object with the correct weights, strata, and primary sampling units for each survey era.

Status: 40 survey years, 1985 through 2024, are published as data releases. BRFSS began in 1984, but CDC no longer serves the 1984 file. Its archived documentation page describes 12,258 records from the first 15 states, and the download link there no longer resolves (checked August 2026), so the collection begins at 1985. brfss_years() reports the currently hosted years, refreshing its manifest at most once a day (refresh = TRUE forces a fresh look).

Installation

Install from CRAN:

install.packages("brfssdata")

Or install the development version from GitHub:

# install.packages("pak")
pak::pak("muntasirmasum/brfssdata")

First estimate

Three lines from install to a design-correct estimate. The first call downloads the year once (about 29 MB, checksum-verified), then everything is served from the local cache.

library(brfssdata)
library(srvyr)

brfss_design(2023, vars = "GENHLTH") |>
  group_by(GENHLTH) |>
  summarize(prop = survey_prop(vartype = "ci"))
#> # A tibble: 6 × 4
#>   GENHLTH    prop prop_low prop_upp
#>     <dbl>   <dbl>    <dbl>    <dbl>
#> 1       1 0.158    0.156    0.161
#> 2       2 0.310    0.307    0.313
#> 3       3 0.335    0.332    0.339
#> 4       4 0.149    0.146    0.151
#> 5       5 0.0444   0.0431   0.0458
#> 6      NA 0.00316  0.00276  0.00362

Weights, strata, and primary sampling units are already applied, with the right weight for the survey era; codes CDC uses for "don't know" and "refused" arrive as NA (that is the na = TRUE default on brfss_design()), so the proportions cover substantive answers. Add labels = TRUE if you want the factor levels spelled out.

Usage

library(brfssdata)

# Which survey years are published?
brfss_years()

# Respondent-level data, only the variables you need
dat <- read_brfss(2019:2023, vars = c("GENHLTH", "PHYSHLTH", "_LLCPWT"))

# The same, with safe categoricals converted to labeled factors
dat <- read_brfss(2023, vars = c("GENHLTH", "SEXVAR"), labels = TRUE)

# One state's rows only, filtered inside the query
dat <- read_brfss(2023, vars = "GENHLTH", states = "TX")

# The value-label codebook itself (1998 onward)
brfss_labels("GENHLTH", years = 2023)

# Everything the catalogs know about a variable, as a card
brfss_codebook("GENHLTH", years = 2023)

# CDC renames variables across years; find the whole family
brfss_crosswalk("_DRNKWK1")

# A survey-design object with era-correct weights, ready for srvyr.
# Codes CDC uses for "don't know" and "refused" are NA by default here,
# so the proportions cover substantive answers; see ?brfss_missing_codes.
library(srvyr)
brfss_design(2023, vars = "GENHLTH") |>
  group_by(GENHLTH) |>
  summarize(prop = survey_prop(vartype = "ci"))

# Where does a variable appear across years?
brfss_vars("smok")

Downloads are verified against published checksums and cached under brfss_cache_dir() (per tools::R_user_dir()). One call to brfss_download(2019:2023) prefetches years plus the variable and label catalogs, after which everything runs offline; brfss_cache_info() and brfss_cache_clear() manage the cache.

The files come from GitHub releases (a survey year from 2011 on is 20 to 45 MB; all 40 years about 737 MB), so a network that blocks github.com, a TLS-intercepting proxy, or an air-gapped compute node blocks the download too. In that case prefetch on an unrestricted machine, copy the cache directory across, and point options(brfssdata.cache_dir = ...) at it; one shared directory serves a whole lab or cluster. The "Caching and offline work" section of the getting-started vignette walks through it.

The 2011 design break

BRFSS added cell-phone-only respondents and switched from post-stratification to raking in 2011; CDC states that estimates from 2011 onward are not directly comparable to earlier years. brfss_design() selects the era-correct weight automatically (_FINALWT before 2011, _LLCPWT after) and refuses to pool years across the boundary unless you opt in with allow_break = TRUE.

Access from Python, SAS, Stata, or anything else

The hosted data files are plain parquet, so no R is required to use them. Every release asset has a stable URL:

import pandas as pd

url = ("https://github.com/muntasirmasum/brfssdata/releases/download/"
       "data-2023/brfss_2023.parquet")
df = pd.read_parquet(url)

The same URLs work in polars, DuckDB (any language), Julia, or Stata's Python bridge. SAS and Stata users who want native files can export any extract from R with haven::write_xpt() (SAS Transport) or haven::write_dta(); the Using the data outside R article walks through both, including the variable-name limits each format imposes on CDC's calculated names.

For survey-weighted analysis outside R, use the design columns shipped in every year: the final weight (_LLCPWT from 2011, _FINALWT before), strata (_STSTR), and PSU (_PSU). The list of available years lives at releases/download/data-meta/manifest.json, and each file has a .sha256 companion for verification.

Getting help

Questions and problems go to the issue tracker at https://github.com/muntasirmasum/brfssdata/issues, where templates cover bugs, data problems, and feature requests. For a bug, include a minimal example and the output of sessionInfo(). For a data problem, give the survey year, the variable name as CDC spells it, and the code or label in question, because the answer usually sits in CDC's codebook for that year. Add the output of brfss_cache_info() whenever a download or a cached file is involved. The function reference and the articles on the package site are the first place to look.

Contributing

Contributions are welcome. CONTRIBUTING.md explains how to set up a development copy, run the offline test suite, and propose changes to the data build. Please note that the brfssdata project is released with a Contributor Code of Conduct. By contributing to this project, you agree to abide by its terms.

Citation

To cite the package, use citation("brfssdata"):

Masum M (2026). brfssdata: Access CDC Behavioral Risk Factor Surveillance System Data. doi:10.32614/CRAN.package.brfssdata https://doi.org/10.32614/CRAN.package.brfssdata. R package version 0.1.0, https://muntasirmasum.github.io/brfssdata/.

Analyses should also cite the underlying survey data (below).

The source code of every release is also archived on Zenodo under the concept DOI 10.5281/zenodo.21972542, which resolves to the latest archived version.

Data source

All data originate from the CDC BRFSS annual survey files, which are in the public domain. Suggested citation for the data:

Centers for Disease Control and Prevention (CDC). Behavioral Risk Factor Surveillance System Survey Data. Atlanta, Georgia: U.S. Department of Health and Human Services, Centers for Disease Control and Prevention, [appropriate year].

The hosted parquet files are derived from CDC's published SAS Transport files with no rows dropped, no columns dropped, and no values recoded. Three transformations are applied: the Windows-1252 text older files carry is re-encoded to UTF-8, blank SAS character fields (SAS's missing value for character data) are stored as nulls rather than empty strings, and an integer year column is added. The handful of variables CDC stored as a number in some years and text in others are written with one storage type across all years, values unchanged, so multi-year reads cannot split one code into 1120 and "1120.0". The processing pipeline lives in data-raw/; every artifact is checksummed, and the package verifies those checksums when it downloads.

License

MIT for the package code. The BRFSS data are a U.S. government work in the public domain.

About

R data package for CDC BRFSS: external parquet hosting, DuckDB lazy queries, srvyr survey designs

Topics

Resources

Code of conduct

Contributing

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages