Skip to content

Repository files navigation

snmpio

An async C++20 library for SNMPv2c and SNMPv3 command generation — GET, GETNEXT, GETBULK, SET and subtree walks — built directly on Asio with no net-snmp dependency. Manager side only.

The domain vocabulary this codebase uses is defined in CONTEXT.md; the decisions that shaped it are in docs/adr/.

Status

Stage 4 of 6. SNMPv2c and SNMPv3 both work end to end over UDP: GET, GETNEXT, GETBULK, SET and Walk, at all three Security Levels. Engine Discovery, time synchronisation and Report routing happen underneath and are never surfaced. authPriv speaks DES, AES-128, and AES-192/256 under both the Blumenthal and the Reeder key extension. Stage 5's automated half is done too: every operation reaches snmpd and both Simulator images in CI, and what stage 5 still needs is a run of the hardware checklist.

Stage Deliverable State
0 CMake skeleton, OID/value types, BER encode/decode + fuzz targets done
1 v2c GET / GETNEXT / GETBULK / SET and Walk over Asio UDP done
2 v3 message framing, USM auth (MD5, SHA-1, SHA-2), password-to-key, key localization done
3 Async engine discovery, time sync, Report handling done
4 Privacy: AES-128, then AES-192/256 under both key extensions, DES behind the legacy provider done
5 Interop matrix vs the Simulator, snmpd, and real vendor gear automated half done; the hardware checklist remains
6 Docs, cancellation semantics, error taxonomy, packaging

Using it

Both snippets below are condensed from examples/, which is a separate CMake project that find_package()s an installed snmpio — so building it is what proves the install rules work, and CI builds it on every push: the examples and the package cannot drift apart. There is one example per operation; the table below says what each shows.

A v2c GET, in the callback form:

namespace net = snmpio::net;
net::IoContext io;
snmpio::Client client(io.get_executor());

snmpio::Target target;
target.endpoint = {net::asio::ip::make_address("127.0.0.1"), 161};

client.asyncGet(target, snmpio::Community{"public"}, {snmpio::Oid{1, 3, 6, 1, 2, 1, 1, 1, 0}},
                [&client](const net::ErrorCode& ec, const snmpio::Response& response) {
                  client.stop();
                  if (ec) return;
                  std::cout << snmpio::toString(response.varbinds.front().val) << "\n";
                });

io.run();

A v3 authPriv Walk, in the coroutine form. There is nothing extra to call first: Engine Discovery and time synchronisation happen underneath, and swapping the Community for Credentials is the whole of the difference between the two versions at this API.

snmpio::Credentials credentials;
credentials.userName = "privsha1aes";
credentials.level = snmpio::SecurityLevel::AuthPriv;
credentials.authProtocol = snmpio::AuthProtocol::Sha1;
credentials.authPassword = password;
credentials.privProtocol = snmpio::PrivProtocol::Aes128;
credentials.privPassword = password;

net::ErrorCode ec;
auto varbinds = co_await client.asyncWalkCollect(
    target, credentials, *snmpio::Oid::parse("1.3.6.1.2.1.1"), {},
    net::asio::redirect_error(net::asio::use_awaitable, ec));

Three things the compiler will not tell you:

  • The Security Level is required, never inferred. A Client that silently downgraded authPriv because the Credentials happened to carry no privacy password would be a security hole, so an authPriv level with PrivProtocol::None is Errc::UnsupportedPrivProtocol rather than an authNoPriv request.
  • io.run() returns only after client.stop(). The Client's receive loop is outstanding work. stop() also fails everything in flight with Errc::ClientStopped; it is deliberately not called from the destructor, because the cleanup runs on the strand and would be scheduled against an object that no longer exists.
  • A Target is an address, not a hostname. Choosing a resolver stays the caller's business (CONTEXT.md), so nothing here will quietly resolve one for you.

The examples

Between them, the six cover every operation, every Security Level, and every completion style. Each fixes its protocols in code rather than parsing them from the command line, so what you copy is the API, not an argument parser; Usm.hpp lists the alternatives.

Example Operation Identity Completion Also shows
get asyncGet v2c Community callback the three error categories in one handler
getnext asyncGetNext v3 noAuthNoPriv use_future the blocking style; several OIDs in one request
getbulk asyncGetBulk v3 authNoPriv, SHA-512 coroutine nonRepeaters and maxRepetitions
set asyncSet v3 authPriv, SHA-256 / AES-128 callback a refusal's error-status and error-index, told apart from every other failure
walk asyncWalk v3 authPriv, SHA-1 / AES-256 (Reeder) coroutine streaming; a row limit as a total cancellation, ending in WalkIncomplete
walk-collect asyncWalkCollect v3 authPriv, SHA-1 / AES-128 coroutine the buffering convenience over walk

To build them and run each against the snmpd the interop script starts, whose users are named after what they carry and whose password is snmpio-interop:

cmake --install build/default --prefix /tmp/prefix
cmake -S examples -B build/examples -DCMAKE_PREFIX_PATH=/tmp/prefix
cmake --build build/examples
cd build/examples
./example-get 127.0.0.1 16161 public
./example-getnext 127.0.0.1 16161 noauth 1.3.6.1.2.1.1.1 1.3.6.1.2.1.1.3 1.3.6.1.2.1.1.5
./example-getbulk 127.0.0.1 16161 authsha512 snmpio-interop
./example-set 127.0.0.1 16161 writer-privsha256aes snmpio-interop ops@example.net
./example-walk 127.0.0.1 16161 privsha1aes256c snmpio-interop 1.3.6.1.2.1.1
./example-walk-collect 127.0.0.1 16161 privsha1aes snmpio-interop 1.3.6.1.2.1.1

Two variations show the failure paths. set as privsha256aes — the same protocols, a user that may only read — prints the Agent's refusal, noAccess (error-status 6), error-index 1. A user on other protocols, such as privsha1aes, gets no refusal at all: its digest does not match, which is a different failure. And walk with a trailing row limit, ... 1.3.6.1.2.1.1 5, stops after five rows with WalkIncomplete.

Errors

Every operation reports failure the Asio way, as a net::ErrorCode, and three categories can land in the same completion handler: the system's for socket faults, snmpio's (Errc) for timeouts and unusable replies, and snmp-agent's (ErrorStatus) for an error-status the Agent itself returned. ec.message() is readable in all three.

Writing a retry policy against those individually means switching over forty-odd enumerators, and getting it wrong the same way every time — a wrong password and a lost datagram both look like failure, and retrying the first is pointless while retrying the second is the whole reason UDP transport is survivable. classify() answers the question that actually matters, across all three categories at once:

ErrorClass Means Do
Ok not a failure carry on
Retriable the Target was silent, unreachable, or busy retry unchanged, with backoff
Configuration something you set is wrong or unacceptable fix the Credentials, the Oid, or the request size — then retry
Fatal nothing you can change will help give up on this request
Unclassified an ErrorCode from a fourth category treat as Fatal
snmpio::Response response;
net::ErrorCode ec;

for (int attempt = 0; attempt < 3; ++attempt) {
  response = co_await client.asyncGet(target, community, oids,
                                      net::asio::redirect_error(net::asio::use_awaitable, ec));
  if (snmpio::classify(ec) != snmpio::ErrorClass::Retriable) break;
  co_await backOff(attempt);
}

if (ec) std::cerr << ec.message() << "\n";  // Configuration, Fatal, or Retriable but exhausted
co_return response;

snmpio::Client already retransmits inside a single request, up to Target::retries, before reporting Errc::Timeout — the loop above is the layer above that, for a Target that stayed unreachable across whole exchanges.

The borderline calls are documented next to the enumeration in include/snmpio/Error.hpp, with the reasoning. Three worth knowing here:

  • Errc::AuthFailed is Configuration, not Retriable. A wrong authentication password fails identically on every retry, so a loop that waits it out never terminates against a Target that is answering perfectly. The same goes for DecryptionFailed and UnknownUserName.
  • Errc::NotInTimeWindow is Retriable, even though the Client already resynchronised once before reporting it. It is what an Agent that rebooted mid-exchange produces, and the next request discovers the new boots/time.
  • ErrorStatus::CommitFailed is Fatal. The Agent is saying it does not know what state the SET left behind; replaying a half-applied SET is the one retry that can do damage.

Building

cmake --preset default      # or: standalone, debug, asan, tsan, tidy, fuzz
cmake --build --preset default
ctest --preset default

Requires a C++20 compiler with coroutine support, CMake 3.24+, OpenSSL 3.0 or newer, and either Boost.Asio 1.77 or newer (the default) or standalone Asio 1.21 or newer. The Asio floor is per-operation cancellation, which the Walk's total/terminal split is built on. CMake enforces it for Boost and for standalone Asio found via its config package; the bare-include-directory fallback has no version to check.

OpenSSL supplies the hashes and HMACs USM needs (ADR-0001). It is required rather than optional: SNMPv3 is the reason this library exists, and a build with USM silently missing would be a trap.

To avoid the Boost dependency, use the standalone preset — standalone Asio is asio on Arch and libasio-dev on Debian/Ubuntu.

Preset What it is
default Boost.Asio, RelWithDebInfo
standalone Standalone Asio, RelWithDebInfo
debug Boost.Asio, Debug
asan Boost.Asio, Debug, address + undefined-behaviour sanitizers
tsan Clang, Boost.Asio, Debug, thread sanitizer
tidy Clang with clang-tidy folded into the build
fuzz Clang, fuzzers on, tests off

Each writes to build/<preset>/. Editors that read CMakePresets.json (VS Code CMake Tools, CLion, Qt Creator) will offer these directly.

Option Default Meaning
SNMP_USE_BOOST_ASIO ON Build against Boost.Asio rather than standalone Asio (ADR-0002)
SNMPIO_BUILD_TESTS on if top-level Build the GoogleTest suite
SNMPIO_BUILD_FUZZERS OFF Build the libFuzzer targets (Clang only)
SNMPIO_SANITIZE OFF Address and undefined-behaviour sanitizers
SNMPIO_TSAN OFF Thread sanitizer (Clang only; not with SNMPIO_SANITIZE)
SNMPIO_WERROR OFF Treat warnings as errors

The Asio choice appears in every public signature, so it is resolved at configure time rather than at first use — a consumer who gets it wrong finds out from CMake instead of from a page of template errors. CI builds both.

Sanitizer and hardening builds

A green sanitizer run should mean the sanitizer looked, so each build is set up to see what it claims to:

  • Asio's memory recycling is off in every sanitizer build — asan, tsan and fuzz. Otherwise Asio hands a freed handler or coroutine frame to a thread-local cache and reuses it, and a use-after-free touches a live block that ASan has no reason to report.
  • tsan knows Asio's synchronisation. TSan does not model std::atomic_thread_fence, which Asio's fenced blocks are built on. Asio 1.38.2 annotates them for TSan; on anything older, the tsan build switches them off and the configure log says so. Use a current Clang: the CI cell is on clang-20, because clang 18 predates the fix for a coroutine race of the compiler's own making (llvm#72006).
  • Standard-library hardening — _GLIBCXX_ASSERTIONS, and libc++'s fast hardening mode where libc++ is used — is on for this repository's own targets in every build, so an out-of-bounds operator[] aborts instead of reading on. It is never exported: a consumer of the installed package builds with whatever flags they choose, and CI's install-and-consume job checks that.

docs/research/snmpio-safety-threats.md cites the source for each: §3.4 for the compiler bug, §4–5 for the rest. The one addition is sizing Asio's recycling cache to zero, which older Asio needs because it lacks one of the note's macros; CMakeLists.txt cites the Asio header at the definition.

CI runs the suite under ASan+UBSan and under TSan on every commit, and runs one interop cell — against snmpd — from an ASan+UBSan build. There is deliberately no MemorySanitizer build, which would need OpenSSL and the standard library instrumented too, and no GCC ASan build alongside Clang's.

Naming

types, files UpperCamelCase
functions, variables lowerCamelCase
private members m_lowerCamelCase
namespaces lower_case

clang-format only reflows code, so the convention is enforced by .clang-tidy's readability-identifier-naming. Public struct fields stay plain lowerCamelCase — m_ marks encapsulated state, and vb.m_name on a POD is only noise.

Static analysis

.clang-tidy enables bugprone, cert, clang-analyzer, concurrency, cppcoreguidelines, misc, modernize, performance, portability and readability with WarningsAsErrors: '*'. Every disabled check carries its reason in the file — a check switched off because it was noisy once is a check that will not catch the real defect later.

cmake --preset default                          # always writes compile_commands.json
clang-tidy -p build/default src/*.cpp tests/*.cpp fuzz/*.cpp

Both clang tools are pinned to 22.1.8 — their output changes between major versions, so an unpinned local install will reformat files CI then rejects. Match it with pip install clang-format==22.1.8 clang-tidy==22.1.8 if your distro ships something else.

Or fold it into the build with -DSNMPIO_CLANG_TIDY=ON. tests/ and fuzz/ carry an overlay relaxing what only applies to library code (GoogleTest's do-while macros, fixture tables, *Oid::parse("1.3.6.1") on a literal).

Each NOLINT carries its reason at the site: make_error_code, whose spelling is fixed by the standard because both error_code types call it unqualified through ADL; the SNMPIO_REGISTER_ERROR_CODE_ENUM macro, which opens a namespace and so cannot be a template; the fixed PRNG seed in the round-trip sweep, which exists precisely to be reproducible; the std::getenv the interop harness reads its Target from, which runs before any test thread does; socket::close(ec) in the interop relay, whose return value only exists under one of the two Asios (ADR-0002); and the curl the misbehaviour suite drives the Simulator's control UI with, which is one form POST that would otherwise be an HTTP client written here.

Interop tests

ctest runs everything against the Scripted Agent, which shares our own reading of the protocol. The interop suite is the half that does not: it talks to an Agent nobody here wrote, over a real socket. There is no Agent in a bare checkout, so those tests skip unless a Target is named. -R Interop selects exactly them, and every one of them needs an Agent — so a run filtered to Interop that reports all green really did reach one.

One command starts any of CI's three Agents in a container — snmpd, simulator or simulator-release — on 127.0.0.1:16161, waits until it answers, and prints the environment the suite needs for it. It is the command CI's interop jobs run, so a failure there reproduces here. It needs a docker on Linux (Podman's docker shim will do) and nothing else — the snmpd readiness check uses host networking, which Docker Desktop does not give. snmpd is built from a pinned Ubuntu image rather than taken from the workstation, whose snmpd is not CI's.

tests/interop/start-agent.sh snmpd > /tmp/snmpio-interop.env \
  && export $(cat /tmp/snmpio-interop.env)        # bash, zsh and fish alike
ctest --preset default -R Interop --output-on-failure
docker rm -f snmpio-interop                       # when done; the next start replaces it anyway

What it prints, for snmpd, is what the harness reads — and what to set by hand against an Agent the script did not start:

SNMPIO_INTEROP_TARGET=127.0.0.1
SNMPIO_INTEROP_PORT=16161                  # omit for 161
SNMPIO_INTEROP_V3_PASSWORD=snmpio-interop  # 8+ characters; omit to skip the v3 half
SNMPIO_INTEROP_V3_USM_REPORTS=1            # the capability flags, below; empty is unset
SNMPIO_INTEROP_V3_KEY_EXTENSIONS=1
SNMPIO_INTEROP_FAULTS=                     # the Simulators' control UI port
SNMPIO_INTEROP_FAULTS_ENGINE_ID=
SNMPIO_INTEROP_BROKEN_GETNEXT=
SNMPIO_INTEROP_NO_TIME_WINDOW_CHECK=
SNMPIO_INTEROP_SET_REFUSAL=noAccess
SNMPIO_INTEROP_WRITER_COMMUNITY=snmpio-writer       # the Writer Credentials, below
SNMPIO_INTEROP_WRITER_NOAUTHNOPRIV=writer-noauth
SNMPIO_INTEROP_WRITER_AUTHNOPRIV=writer-authsha256
SNMPIO_INTEROP_WRITER_AUTHPRIV=writer-privsha256aes

SNMPIO_INTEROP_TARGET is the address the Agent answers at and SNMPIO_INTEROP_PORT the port, which defaults to 161. SNMPIO_INTEROP_V3_PASSWORD is the password every v3 interop user carries, and gates the v3 tests: the Agent has to be running the configuration tests/interop/snmpd-conf.sh or tests/interop/fault-agent-auth.sh prints, and a switch on the bench is not, so an unset variable skips rather than fails. One value configures the Agent and drives the suite, which is why it is not written down twice: the script takes it from SNMPIO_INTEROP_V3_PASSWORD if that is set, and uses snmpio-interop if not.

Two ways in

The above is the first: the Agent is ours to configure, so the tests know the users by name and walk the whole matrix. It is what CI's three Agents run, and what the script starts.

The second is for an Agent that is not ours — a switch on the bench, whose users and Community someone chose years ago and will not be changing for us. Name the one user it has and what that user carries, and the Community it answers to, and the tests address those instead of the convention:

export SNMPIO_INTEROP_TARGET=10.0.0.7
export SNMPIO_INTEROP_COMMUNITY=bench-ro           # the v2c Community; omit for public
export SNMPIO_INTEROP_V3_USER=netops-legacy        # what the user is actually called
export SNMPIO_INTEROP_V3_AUTH=sha256               # none, md5, sha1, sha224, sha256, sha384, sha512
export SNMPIO_INTEROP_V3_PRIV=aes                  # none, des, aes, aes192, aes256, aes192c, aes256c
export SNMPIO_INTEROP_V3_PASSWORD=bench-secret-123 # both secrets; omit only at noAuthNoPriv
ctest --preset default -R Interop --output-on-failure

That answers whether the run passed. To record it in the pre-release checklist, run the test binary instead, as that section shows, because ctest splits the run summary into pieces.

One user is enough to be useful, because a Target typically has exactly one. The matrix test then covers the single pair that user can serve, and the Key Extension test skips, since it needs four users of its own. The other two v3 tests in that file — Engine Discovery and the wrong-password Report — run against the named user rather than against the conventional one. GETNEXT, GETBULK, the Walks and the refused SET run as that user's pair in place of one pair per Security Level, and the GETBULK and Walk passes across every privacy protocol skip, since they need a user per cipher.

The v2c half has the same two ways in. It sends the Community public, which is what every Agent we configure answers to, unless SNMPIO_INTEROP_COMMUNITY names another. A Community is the whole of v2c's secret, so a named one is sent but never printed: the summary says v2c/named community.

Naming a user or a Community the Target does not have fails the suite, the same as any other variable that is set but unusable: a skip there would report green for Credentials nobody ever reached. A wrong Community fails as a timeout, because an Agent drops a v2c request with the wrong Community without answering. The suite also fails on a protocol name that is not one of the words above; a privacy protocol with no authentication protocol beside it — USM derives the privacy key with the authentication protocol's hash, so there is no privacy without one; an authenticating user with no password to authenticate with; and SNMPIO_INTEROP_V3_AUTH or _PRIV set with no user for them to describe.

One password, used as both the authentication and the privacy secret. A Target whose user carries two different ones cannot be addressed this way yet.

Verifying this way in needs no hardware. The script's snmpd also carries netops-legacy, a user deliberately named after nothing it holds, and netops-ro, a Community other than public. That is what a Target somebody else configured looks like:

tests/interop/start-agent.sh snmpd > /tmp/snmpio-interop.env \
  && export $(cat /tmp/snmpio-interop.env)
export SNMPIO_INTEROP_COMMUNITY=netops-ro SNMPIO_INTEROP_V3_USER=netops-legacy
export SNMPIO_INTEROP_V3_AUTH=sha256 SNMPIO_INTEROP_V3_PRIV=aes
ctest --preset default -R 'Interop(V2c|V3)' --output-on-failure

CI runs exactly this, as a second pass over the same snmpd, so the way in for a Target we did not configure is not merely documented.

Everything else the harness reads

An address, not a hostname. A Target is built from an endpoint so that choosing a resolver stays the caller's business (stage 1), and the harness is a caller like any other — so a hostname is rejected outright rather than quietly resolved. A variable that is set but unusable fails the suite; only an unset one skips it, since a typo that skipped would report green for an Agent it never reached.

The variables below say what the Agent at that Target can do, and each gates the tests that would otherwise be asserting on the Agent rather than on this library. Each is set by whoever starts the Agent, because nothing on the wire announces it — for CI's three, that is the script, which is the one place their flags are said:

Variable Set it when the Agent
SNMPIO_INTEROP_V3_KEY_EXTENSIONS serves the privsha1aes192/256(c) users — AES-192/256 under both schemes
SNMPIO_INTEROP_V3_USM_REPORTS answers a bad digest with a usmStats Report, which RFC 3414 leaves optional
SNMPIO_INTEROP_FAULTS can be told to misbehave — the port the Simulator's control UI is on
SNMPIO_INTEROP_FAULTS_ENGINE_ID offers the engineIDChange fault, which not every Simulator build does
SNMPIO_INTEROP_BROKEN_GETNEXT answers a GETNEXT carrying several Varbinds from the wrong requested OIDs (snmp-fault-agent#11)
SNMPIO_INTEROP_NO_TIME_WINDOW_CHECK skips RFC 3414 section 3.2 step 7a, answering a request outside its Time Window with a Response rather than the notInTimeWindows Report (snmp-fault-agent#18)
SNMPIO_INTEROP_SET_REFUSAL is known to refuse a read-only SET with one error-status — its RFC 3416 name, such as noAccess or notWritable

With SNMPIO_INTEROP_BROKEN_GETNEXT set, GETNEXT sends one Varbind per request instead of several, and its summary rows say so; tests/InteropOperations.hpp says why it names a defect rather than an ability. With SNMPIO_INTEROP_SET_REFUSAL unset, the refused SET is held only to being refused, and its rows say that too. With it set to readOnly, which RFC 3416 says an SNMPv2 entity never sends (snmp-fault-agent#10), the refusal is asserted all the same, and its rows say the Agent is non-compliant.

Writer Credentials

The SET that lands writes sysContact.0, reads it back and restores it, so it changes the Target while it runs. It uses only the Writer Credentials the run names: a v2c Community, and one v3 user per Security Level, each carrying that level's representative pair (SHA-256, then AES-128) and SNMPIO_INTEROP_V3_PASSWORD. On snmpd they are separate from every user and Community the rest of the suite reads with, which all stay read-only.

Variable Names the writer
SNMPIO_INTEROP_WRITER_COMMUNITY over v2c, never printed: the summary says v2c/writer community
SNMPIO_INTEROP_WRITER_NOAUTHNOPRIV at noAuthNoPriv
SNMPIO_INTEROP_WRITER_AUTHNOPRIV at authNoPriv, on SHA-256
SNMPIO_INTEROP_WRITER_AUTHPRIV at authPriv, on SHA-256 with AES-128

With one unset, its write skips, and the run summary lists it as skip, not ok. The script names all four for the Agents it starts, and nothing names them for a Target it did not, so a run against someone else's equipment writes nothing unless you choose to. On snmpd the writers reach sysContact and nothing else, and its configuration leaves sysContact.0 unset, because net-snmp makes an object read-only when its configuration sets it. The Simulator has no per-user access control, so there the writers are ordinary users, and sysContact.0 is the one writable entry its values configuration serves.

The SET that is refused needs no writer and changes nothing: it writes sysDescr.0's own value back with the Credentials every other test reads with. So it runs everywhere, hardware included. A refusal whose error-index names some Varbind other than the one sent is noted on its row, not failed (snmp-fault-agent#12); no flag gates it, and the script says which Agent does it.

CI runs three Agents, one job each, and between them they cover every v3 case above. Neither gate is a Security Level being negotiated: the Simulator infers the level from which protocols a user carries, while this library requires it explicitly, and that divergence is deliberate on both sides — a Client that silently downgraded authPriv would have a security hole, where a test Agent that accepts what arrives is merely convenient (ADR-0006).

Every Agent the script starts is pinned — snmpd by tests/interop/snmpd.Dockerfile, the two Simulator images by digest — and there are two Simulator images on purpose. The script says why for both, once.

The v3 users are a convention the tests share with the two configuration generators the script runs, snmpd-conf.sh and fault-agent-auth.sh — noauth, auth<hash> per authentication protocol, and priv<hash><cipher> per pair — because they are ours to create; the Simulator's own example configuration names them otherwise, which is why ours is mounted over it. What the Simulator serves is ours too: fault-agent-values.sh writes the values.json mounted beside it, with fifty instances under interfaces for the Walks and the writable and read-only entries a SET needs. The matrix is MD5, SHA-1 and the four SHA-2 hashes, each of them alone at authNoPriv and again over DES and AES-128 at authPriv. AES-192/256 under both Key Extensions are four more users, and both Agents carry them, so each scheme is read on every commit by an implementation that is not ours as well as by the Simulator, which is. The snmpd Debian, Ubuntu and Arch ship — the script's included — speaks all four (ADR-0006 records which versions). One built without net-snmp's --enable-blumenthal-aes does not, and fails those four rather than skipping them, so leave SNMPIO_INTEROP_V3_KEY_EXTENSIONS unset against it. All four are paired with SHA-1 on purpose, for the reason CoversBothKeyExtensions in tests/TestInteropV3.cpp gives.

What the suite proves: a v2c GET of sysDescr.0, which it prints because no two Agents say the same thing; the eighteen v3 pairs above and the four Key Extension ones; GETNEXT and GETBULK over v2c and at each Security Level on SHA-256 with AES-128, and GETBULK again under every privacy protocol. Those two assert successor semantics without pinning any MIB contents: GETNEXT of system is sysDescr.0, which every Agent here already has; a GETBULK's column from system is its column from sysDescr.0 one row late, strictly increasing until it reaches endOfMibView; and GETNEXT of sysDescr.0 is where that second column starts. And a Walk of interfaces in GETNEXT mode and in GETBULK mode, each streaming and collecting, over v2c, at each Security Level and under every privacy protocol. The Subtree is several batches long on the three Agents CI starts, and on most Targets beside them; on one with too few interfaces the Walks fail rather than pass on a single batch. The Walks assert structure alone — every OID inside the Subtree, strictly increasing, a clean end, more than one batch — and are held to one OID list, so the Agent supplies the expected answer; values are not compared, since counters move between Walks. And SET, over v2c and at each Security Level: a write that lands, read back with a GET and restored, and a write the Agent refuses, asserting the exact error-status its capability flag names and that the value is unchanged afterwards. It also proves that Engine Discovery costs the extra round trips exactly once, counted off the wire by a relay between Client and Agent, since the API deliberately never surfaces it; and that a wrong password comes back as the Report the Engine sent rather than as a timeout.

The misbehaviour suite

SNMPIO_INTEROP_FAULTS is the port of the Simulator's web UI, on the Target's own address — the same endpoint a browser opens, since what the UI does is post a form — and turns on the half of the suite the other Agents cannot run. snmpd and a switch on the bench are correct, and a correct Agent never produces any of these conditions, which is the whole reason the Simulator is CI's primary target (ADR-0006):

What the Agent does What this library has to do
restarts, so its engine boots jump past the pair we cached resynchronise from the Report and complete the request
reports a lower boots count than the one we hold refuse it rather than cache it (RFC 3414 §2.2.3) — being walked backwards is a replay window
steps its clock back inside one boot refuse that too: it is the same comparison and the commoner case, since it needs only NTP
answers a Walk's GETBULK with tooBig ask for fewer repetitions, and finish the Walk once it fits
echoes the requested OID straight back fail the Walk with Errc::NonIncreasingOid instead of asking for ever (ADR-0004)
truncates the Response mid-message drop it and leave the request outstanding until it times out
comes back under a different engineID re-discover it and re-derive both keys against the new one, which every key is localized to

The last one is the one worth stating as a rule: an undecodable datagram says nothing about whether the Response is still coming, so Errc::Timeout is the only honest answer. Surfacing the decode error would let anyone able to send this Client one junk datagram end a request it had no part in.

Each case has a counterpart against the Scripted Agent, for the reason tests/TestInteropFaults.cpp opens with.

Pre-release hardware checklist

CI does not run this, and never will. Everything under Interop tests runs on every push against three Agents CI starts for itself. The rows below are real equipment on a bench, run by hand before a release. A row's date is the last time someone did that; the library has not been checked against that hardware since then.

Each row is there because the automated matrix leaves a gap that only that equipment can close. ADR-0006 assigns those gaps. Hardware that would close no gap gets no row, because nobody would ever re-run it. That is why iLO 5 and Meinberg NTP servers, which ADR-0006 first named in the fleet, have none (its 2026-09-29 amendment).

Row Gap it closes Device and firmware Protocols exercised Last run
Cisco switch Reeder (aes192c/aes256c) against a vendor's Engine. snmpd and the Simulator both check the Reeder path on every commit, and snmpd's reading of the expired draft is independent of ours, but neither is the reading the vendors who actually ship Reeder made — — —
HPE iLO 6 The widest protocol range of any vendor Agent on the bench, SHA-2 and AES-192/256 included. snmpd and the Simulator read all of it but 3DES on every commit, and neither is a vendor's Engine. Its 3DES has to wait, because the library does not speak 3DES and no stage carries it yet — — —

A dash means the run has not happened. Nothing goes into a row that a run did not print. Device and firmware is whatever sysDescr.0 says, verbatim. If a vendor leaves the firmware version out of it, the row leaves it out too, rather than adding it from memory.

Running it against a switch

A switch on the bench carries users someone else named, so it is always addressed the second way in: name its Community, and the one user and the protocols it carries. Do not set up the noauth/auth<hash>/priv<hash><cipher> convention that CI's Agents use.

Of the capability variables, leave SNMPIO_INTEROP_FAULTS and _FAULTS_ENGINE_ID unset, because a correct Agent cannot misbehave on request, and SNMPIO_INTEROP_BROKEN_GETNEXT and _NO_TIME_WINDOW_CHECK unset. Set SNMPIO_INTEROP_SET_REFUSAL only once you know which error-status the Target refuses a read-only SET with; unset, the refused SET still runs, held only to being refused. Name no Writer Credentials unless the Target is yours to write sysContact.0 on; its write rows then say skip. Set SNMPIO_INTEROP_V3_USM_REPORTS only if the Target answers a bad digest with a usmStats Report. That Report is optional behaviour RFC 3414 allows a correct Agent, and a Target that sends it can prove the wrong-password test. SNMPIO_INTEROP_V3_KEY_EXTENSIONS does nothing here, because that test needs four users and a named run has one.

ctest is fine for checking whether the run passed. To fill in a row, run the test binary instead. ctest starts every test as a separate process, so the summary comes out in pieces, one per test, and it only shows a passing test's output under -V. The binary runs the whole interop suite in one process and ends with one summary:

export SNMPIO_INTEROP_TARGET=10.0.0.7 SNMPIO_INTEROP_COMMUNITY=bench-ro
export SNMPIO_INTEROP_V3_USER=netops-legacy SNMPIO_INTEROP_V3_AUTH=sha1
export SNMPIO_INTEROP_V3_PRIV=aes256c SNMPIO_INTEROP_V3_PASSWORD=bench-secret-123
./build/default/tests/snmpio_tests --gtest_filter='Interop*'
== snmpio interop run summary ==
Device/firmware: <the Target's sysDescr.0, verbatim>
Date: <YYYY-MM-DD>

Protocols exercised:
  ok   v2c/named community
  ok   authPriv/sha1/aes256c
  skip noAuthNoPriv  -- run named one user: netops-legacy carrying authPriv/sha1/aes256c
  ...

Every column in the table that a run fills comes from one line of that output. Copy the Device/firmware: line into Device and firmware, the ok lines into Protocols exercised, and the Date: line into Last run. skip lines stay out of the row, because they are protocol pairs the run never reached. A fail line means there is no row to fill. That run found a bug, not a release.

A Target whose users carry more than one protocol pair needs one run per user. Record one row per run, each copied from its own summary, rather than merging several runs from memory into one row.

Fuzzing

cmake --preset fuzz
cmake --build --preset fuzz
mkdir -p .fuzz-work
./build/fuzz/fuzz/FuzzV2cMessage .fuzz-work fuzz/corpus

The first directory is where libFuzzer writes what it finds; fuzz/corpus is passed read-only so the curated seeds stay curated. To replay the seeds alone, as CI does before it fuzzes:

./build/fuzz/fuzz/FuzzV2cMessage fuzz/corpus -runs=0

The fuzzers' asserts are their oracles, so they are live in every fuzz build: the build type defines NDEBUG, and the fuzz targets undefine it. The targets are also built under ASan+UBSan with Asio's recycling off and standard-library hardening on, as the sanitizer builds are.

Five targets, each asserting a round-trip identity rather than merely "does not crash":

  • FuzzBerValue — anything the value decoder accepts must re-encode and decode back identically.
  • FuzzBerVarbindList — the same, over the nesting path: scope entry, length patching, and the trailing-data checks a flat value never reaches.
  • FuzzOidText — the dotted-decimal parser, which is where untrusted text enters the OID type.
  • FuzzV2cMessage — the whole datagram: framing, version, community and the PDU inside them. This is the surface a hostile Agent actually reaches.
  • FuzzV3Message — the v3 datagram, verifyAuth over whatever it decodes to, and decryptScopedPdu over an encryptedPDU. The digest's offset is derived from attacker-controlled length fields and is then used to index the datagram, which is exactly the shape of bug a fuzzer under ASan finds and review does not; decryption then hands a buffer of noise to the BER decoder, which is the same shape one layer down.

What stage 4 contains

  • snmpio::PrivProtocol and two more Credentials fields — the privacy protocol and its own secret. There is no second hash to name: USM derives the privacy key with the authentication protocol's hash. authPriv with no privacy protocol fails with Errc::UnsupportedPrivProtocol rather than being sent in the clear.
  • DES-CBC (RFC 3414 section 8) and AES-CFB128 at 128, 192 and 256 bits (RFC 3826 and the Blumenthal draft). ADR-0005 is why DES is here at all; it lives in OpenSSL 3.x's legacy provider, which is loaded lazily on first use, so a build without it loses that operation and not the library.
  • Both key extensions, because the Localized Key is shorter than an AES-192/256 key whenever the hash is. Aes192/Aes256 are Blumenthal — append the hash of the key so far — and Aes192C/Aes256C are Reeder, which runs the key back through password-to-key and localizes it again. They are separate enumerators rather than a protocol plus a flag, so "AES-192, extension unspecified" is a state that cannot be written down (CONTEXT.md: never inferred, never guessed).
  • Encryption inside the message layer, not beside it: encodeV3Message encrypts the ScopedPDU and writes the salt it chose into msgPrivacyParameters, then computes the digest over the finished message — so authentication covers the ciphertext, and the two are done in the order RFC 3414 section 3.2 checks them in.
  • decodeV3Message stops at the ciphertext and decryptScopedPdu opens it, because the key is a property of the request this answers and finding that request needs the msgID the decode produced. A reply that will not decrypt is dropped exactly like one whose digest is wrong: the request stays outstanding and its retransmission timer keeps running.
  • The privacy key is cached beside the authentication one, on (engineID, hash, secret, privacy protocol). The protocol is part of the cache key because it decides how far the derivation is extended — and under Reeder that extension is a second megabyte hash.

"DES behind an opt-in", as the stage was first written, is the legacy provider rather than a build flag: ADR-0005 rules a build flag out, so the opt-in is naming PrivProtocol::Des on an OpenSSL that has the provider. Nothing else changes shape for it.

Not in stage 4: 3DES, which ADR-0005 names alongside DES and no stage yet carries -- an open gap against that ADR rather than a decision against it -- and IDEA, which ADR-0005 excludes outright.

What stage 3 contains

  • Client's six operations again, taking Credentials where the v2c ones take a Community. Same completion signatures, same three error categories; the type of that one argument is the whole of the difference at the call site.
  • Engine Discovery, RFC 3414 section 4, as ordinary async work on the Client's existing strand — which is the entire point of ADR-0001, since net-snmp doing this synchronously inside its send path is why this library exists. Phase one learns the engineID; phase two, needed only when authenticating, learns the boots/time pair. Requests arriving while a discovery is in flight queue behind it rather than each probing separately.
  • Caches, owned by the Client because ADR-0003 says there is nowhere else, and keyed the way that ADR requires: the Authoritative Engine on its engineID, with a separate endpoint→engineID index, and the Localized Key on (engineID, protocol, secret). One Engine reachable at two Targets is therefore one cache entry and one megabyte-hash derivation, which is the whole reason there is no session type.
  • Report routing. The six usmStats counters map to outcomes in one table: notInTimeWindows and unknownEngineIDs resynchronise and retry exactly once, the rest fail the request with an ErrorCode naming which. No Report ever reaches a completion handler.
  • The Time Window, 150 seconds, checked against the cached pair projected forward by the local clock rather than against a raw cached number.
  • Discovery outlives the request that started it (ADR-0003 again): it runs detached on the Client's strand, so cancelling whichever request happened to arrive first does not cancel what every other request is queued behind.

A Response that fails to decode, to authenticate, or to be timely is dropped, not failed — the request stays outstanding and its retransmission timer keeps running. UDP is spoofable and the msgID is guessable, so the alternative is a library whose requests anyone on the path can cancel.

Reports are the exception, and cannot not be. The four counters worth hearing about are exactly the ones an Engine cannot sign — it does not know the user, or the key, or the engineID it was addressed by — so refusing an unauthenticated Report would turn "wrong password" into "timed out". They are accepted against an outstanding msgID from the address we sent to, which is the same bar a spoofed v2c Response clears. What an unauthenticated Report can never do is change cached state: resynchronising the boots/time pair requires a Report whose digest verified, and anything else that asks us to resynchronise gets a full re-discovery instead, whose own answer is authenticated.

Outstanding requests are keyed on the Message ID rather than the PDU's request-id, as CONTEXT.md requires: a message that cannot be opened must still be attributable to the request that sent it.

What stage 2 contains

  • snmpio::AuthProtocol / snmpio::SecurityLevel / snmpio::Credentials — the USM user, the level they authenticate at, and the protocol behind it. MD5 and SHA-1 are present on purpose (ADR-0005). Security Level is valued as its msgFlags bits, the way PduType is valued as its BER tag.
  • snmpio::passwordToKey / snmpio::localizeKey — RFC 3414 appendix A.2's megabyte expansion, and the hash that binds the resulting Master Key to one Engine. Both are checked against the RFC's own MD5 and SHA-1 vectors. Neither caches: a cache without an owner is a leak, and stage 3 owns the per-(Credentials, engineID) one.
  • snmpio::V3Header / snmpio::UsmParameters / snmpio::ScopedPdu — RFC 3412's message framing, RFC 3414's security parameters inside their OCTET STRING, and the PDU with the context it is interpreted in.
  • encodeV3Message / verifyAuth — the digest computed over the finished message with its own field blanked, and checked in constant time on the way back. Each layer is encoded into its own buffer and then wrapped, so that a sequence length widening past 127 Octets cannot silently move the digest; the long-user-name test is the one that fails if that changes.

Not in stage 2: privacy, which arrived with stage 4 and put the encryption either side of the digest in encodeV3Message and decryptScopedPdu; and timeliness, which needs the cached engine state discovery produces and so arrived with stage 3.

What stage 1 contains

  • snmpio::Client — the Command Generator. It owns the sockets, the outstanding-request table and the strand everything internal runs on; there is no session type (ADR-0003). asyncGet, asyncGetNext, asyncGetBulk, asyncSet, asyncWalk and asyncWalkCollect all take an Asio completion token and report failure as an error_code.
  • snmpio::Target / snmpio::Community — a transport endpoint with its timeout and retry count, and the string that authorizes a v2c request. Neither knows about the other.
  • snmpio::Pdu / snmpio::PduType — the RFC 3416 PDUs, plus the SNMPv2c message framing around them. GETBULK's non-repeaters and max-repetitions are the error-status and error-index slots, named by accessors rather than duplicated into a second struct.
  • snmpio::ErrorStatus — the Agent's own error-status, in a category of its own so that its numbering stays the RFC's. A tooBig from an Agent is never confused with one of our faults.

Three failure channels reach a completion handler, and they stay distinguishable: the system category for socket faults, snmpio for timeouts and malformed Responses, snmp-agent for an error-status the Agent returned.

Walks stream by default and collect on request (ADR-0004). Both reject a non-increasing OID, both stop at the Subtree boundary, and both degrade max-repetitions when an Agent answers tooBig. Cancellation is split as the ADR requires: total stops at a batch boundary and reports Errc::WalkIncomplete alongside what was already delivered, terminal drops everything with operation_aborted.

A single request reads both signals the same way whichever wait it is in -- awaiting a reply, between retransmissions, or queued behind an Engine Discovery: terminal drops it at once, total stops it cleanly but still takes a reply already on its way, and either completes with operation_aborted rather than Errc::Timeout.

Not in stage 1, and deliberately: SNMPv3 in any form (stages 2 and 3), and hostname resolution — a Target is built from an asio::ip::udp::endpoint, so resolving is the caller's choice of resolver rather than a policy this library picks, in this stage or any later one.

What stage 0 contains

  • snmpio::Oid — a permissive OID type with lexicographic ordering and subtree-prefix testing. It will hold sequences X.690 cannot encode, because it also has to represent what a misbehaving Agent sent; isEncodable() is the separate question.
  • snmpio::Value / snmpio::Varbind — a variant over the RFC 2578 application types and the three RFC 3416 exception markers. Counter32, Gauge32 and TimeTicks are distinct types, not aliases of uint32_t.
  • snmpio::ber::Reader / snmpio::ber::Writer — a non-throwing codec with sticky errors, so a decoder is a straight run of reads with one check at the end. The reader clamps every read to the element it is inside; the writer patches sequence lengths in place.
  • snmpio::Errc — the codec error taxonomy, registered with whichever error_code the Asio choice selected.

The codec is deliberately lenient where leniency is unambiguous and strict where it is not: redundant integer sign padding and non-minimal long-form lengths are accepted, because agents emit them and there is only one thing they can mean; indefinite lengths, high-tag-number form, oversized sub-identifiers and non-minimal OID sub-identifiers are rejected.

What it deliberately does not contain

No MIB parsing, in this stage or any later one. GET, SET and WALK operate on numeric OIDs, and nothing in RFC 3411/3412/3414/3416/3417 requires MIB knowledge. If symbolic names are ever wanted they belong in a separate optional target consuming pre-compiled MIB data.

Contributing

Issues and specs live in GitHub Issues. Triage uses five labels — needs-triage, needs-info, ready-for-agent, ready-for-human, wontfix.

License

Apache-2.0 — see LICENSE and NOTICE, and ADR-0007 for why. Same licence as snmp-fault-agent, the simulator CI tests against.

There are no per-file copyright headers, deliberately: copyright is automatic, and a boilerplate block on every file is upkeep that buys nothing.

About

An async C++20 SNMPv2c/SNMPv3 command generator built directly on Asio — GET, GETNEXT, GETBULK, SET and streaming subtree walks, with no net-snmp dependency.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages