brfssdata gives R users direct, reproducible access to annual microdata from the CDC Behavioral Risk Factor Surveillance System (BRFSS), the largest continuously conducted health telephone survey in the world.
Survey years are processed once into compact parquet files and hosted as GitHub release assets. The package downloads each requested year a single time into a local cache, queries it through DuckDB (so pulling a handful of variables from a 300-plus column survey is fast), and hands back either a tibble or a ready-made srvyr survey-design object with the correct weights, strata, and primary sampling units for each survey era.
Status: 40 survey years, 1985 through 2024, are published as data releases. BRFSS began in 1984, but CDC no longer serves the 1984 file. Its archived documentation page describes 12,258 records from the first 15 states, and the download link there no longer resolves (checked August 2026), so the collection begins at 1985.
brfss_years()reports the currently hosted years, refreshing its manifest at most once a day (refresh = TRUEforces a fresh look).
Install from CRAN:
install.packages("brfssdata")Or install the development version from GitHub:
# install.packages("pak")
pak::pak("muntasirmasum/brfssdata")Three lines from install to a design-correct estimate. The first call downloads the year once (about 29 MB, checksum-verified), then everything is served from the local cache.
library(brfssdata)
library(srvyr)
brfss_design(2023, vars = "GENHLTH") |>
group_by(GENHLTH) |>
summarize(prop = survey_prop(vartype = "ci"))
#> # A tibble: 6 × 4
#> GENHLTH prop prop_low prop_upp
#> <dbl> <dbl> <dbl> <dbl>
#> 1 1 0.158 0.156 0.161
#> 2 2 0.310 0.307 0.313
#> 3 3 0.335 0.332 0.339
#> 4 4 0.149 0.146 0.151
#> 5 5 0.0444 0.0431 0.0458
#> 6 NA 0.00316 0.00276 0.00362Weights, strata, and primary sampling units are already applied, with
the right weight for the survey era; codes CDC uses for "don't know"
and "refused" arrive as NA (that is the na = TRUE default on
brfss_design()), so the proportions cover substantive answers. Add
labels = TRUE if you want the factor levels spelled out.
library(brfssdata)
# Which survey years are published?
brfss_years()
# Respondent-level data, only the variables you need
dat <- read_brfss(2019:2023, vars = c("GENHLTH", "PHYSHLTH", "_LLCPWT"))
# The same, with safe categoricals converted to labeled factors
dat <- read_brfss(2023, vars = c("GENHLTH", "SEXVAR"), labels = TRUE)
# One state's rows only, filtered inside the query
dat <- read_brfss(2023, vars = "GENHLTH", states = "TX")
# The value-label codebook itself (1998 onward)
brfss_labels("GENHLTH", years = 2023)
# Everything the catalogs know about a variable, as a card
brfss_codebook("GENHLTH", years = 2023)
# CDC renames variables across years; find the whole family
brfss_crosswalk("_DRNKWK1")
# A survey-design object with era-correct weights, ready for srvyr.
# Codes CDC uses for "don't know" and "refused" are NA by default here,
# so the proportions cover substantive answers; see ?brfss_missing_codes.
library(srvyr)
brfss_design(2023, vars = "GENHLTH") |>
group_by(GENHLTH) |>
summarize(prop = survey_prop(vartype = "ci"))
# Where does a variable appear across years?
brfss_vars("smok")Downloads are verified against published checksums and cached under
brfss_cache_dir() (per tools::R_user_dir()). One call to
brfss_download(2019:2023) prefetches years plus the variable and label
catalogs, after which everything runs offline; brfss_cache_info() and
brfss_cache_clear() manage the cache.
The files come from GitHub releases (a survey year from 2011 on is 20 to
45 MB; all 40 years about 737 MB), so a network that blocks github.com,
a TLS-intercepting proxy, or an air-gapped compute node blocks the
download too. In that case prefetch on an unrestricted machine, copy
the cache directory across, and point
options(brfssdata.cache_dir = ...) at it; one shared directory serves
a whole lab or cluster. The "Caching and offline work" section of the
getting-started
vignette
walks through it.
BRFSS added cell-phone-only respondents and switched from
post-stratification to raking in 2011; CDC states that estimates from
2011 onward are not directly comparable to earlier years. brfss_design()
selects the era-correct weight automatically (_FINALWT before 2011,
_LLCPWT after) and refuses to pool years across the boundary unless you
opt in with allow_break = TRUE.
The hosted data files are plain parquet, so no R is required to use them. Every release asset has a stable URL:
import pandas as pd
url = ("https://github.com/muntasirmasum/brfssdata/releases/download/"
"data-2023/brfss_2023.parquet")
df = pd.read_parquet(url)The same URLs work in polars, DuckDB (any language), Julia, or Stata's
Python bridge. SAS and Stata users who want native files can export any
extract from R with haven::write_xpt() (SAS Transport) or
haven::write_dta(); the Using the data outside
R
article walks through both, including the variable-name limits each
format imposes on CDC's calculated names.
For survey-weighted analysis outside R, use the design columns shipped in
every year: the final weight (_LLCPWT from 2011, _FINALWT before),
strata (_STSTR), and PSU (_PSU). The list of available years lives at
releases/download/data-meta/manifest.json, and each file has a
.sha256 companion for verification.
Questions and problems go to the issue tracker at
https://github.com/muntasirmasum/brfssdata/issues, where templates
cover bugs, data problems, and feature requests. For a bug, include a
minimal example and the output of sessionInfo(). For a data problem,
give the survey year, the variable name as CDC spells it, and the code
or label in question, because the answer usually sits in CDC's codebook
for that year. Add the output of brfss_cache_info() whenever a
download or a cached file is involved. The function reference and the
articles on the package site
are the first place to look.
Contributions are welcome. CONTRIBUTING.md explains how to set up a development copy, run the offline test suite, and propose changes to the data build. Please note that the brfssdata project is released with a Contributor Code of Conduct. By contributing to this project, you agree to abide by its terms.
To cite the package, use citation("brfssdata"):
Masum M (2026). brfssdata: Access CDC Behavioral Risk Factor Surveillance System Data. doi:10.32614/CRAN.package.brfssdata https://doi.org/10.32614/CRAN.package.brfssdata. R package version 0.1.0, https://muntasirmasum.github.io/brfssdata/.
Analyses should also cite the underlying survey data (below).
The source code of every release is also archived on Zenodo under the concept DOI 10.5281/zenodo.21972542, which resolves to the latest archived version.
All data originate from the CDC BRFSS annual survey files, which are in the public domain. Suggested citation for the data:
Centers for Disease Control and Prevention (CDC). Behavioral Risk Factor Surveillance System Survey Data. Atlanta, Georgia: U.S. Department of Health and Human Services, Centers for Disease Control and Prevention, [appropriate year].
The hosted parquet files are derived from CDC's published SAS Transport
files with no rows dropped, no columns dropped, and no values recoded.
Three transformations are applied: the Windows-1252 text older files
carry is re-encoded to UTF-8, blank SAS character fields (SAS's missing
value for character data) are stored as nulls rather than empty
strings, and an integer year column is added. The handful of
variables CDC stored as a number in some years and text in others are
written with one storage type across all years, values unchanged, so
multi-year reads cannot split one code into 1120 and "1120.0". The
processing pipeline lives in
data-raw/;
every artifact is checksummed, and the package verifies those checksums
when it downloads.
MIT for the package code. The BRFSS data are a U.S. government work in the public domain.