Skip to content

feat(metrics): opt-in secure gateway metrics endpoint #3664

Description

@gmenher

User Story

As an OpenShell platform operator, I want to opt in to a gateway metrics endpoint that uses verified HTTPS and an authenticated monitoring client, so that I can collect operational metrics under a defined transport and trust contract without disabling certificate verification.

Problem Statement

The Helm chart enables a dedicated gateway metrics listener on port 9090 by default and exposes it through the OpenShell Service as the named metrics port.

The gateway serves GET /metrics directly over a raw TCP listener. The endpoint does not perform a TLS handshake and does not authenticate or authorize the caller. A workload that can reach the Service can read the metrics over plaintext HTTP.

There is currently no metrics TLS configuration, client-authentication configuration, or secure scrape configuration in the chart. Network reachability can be constrained by deployment policy, but the endpoint itself does not provide transport confidentiality, server identity verification, or caller authentication.

Impact / Why This Matters

The currently emitted metrics are low-sensitivity operational signals: gRPC method and status counts, request latencies, normalized HTTP paths, database readiness, and gateway-interceptor names and binding identifiers. They do not contain secrets, identities, or sandbox names. Plaintext scraping on a dedicated ClusterIP port is therefore a valid default Prometheus workflow.

Some operators nevertheless require encrypted transport, verified server identity, and authenticated scrape access for every observability endpoint. Today they can disable the metrics endpoint or rely on deployment-specific network policy, but neither option provides a supported secure scrape contract. An opt-in TLS and mTLS mode gives those operators a portable solution while preserving the default behavior for existing deployments.

Proposed Design

OpenShell should add an optional secure metrics mode. Plaintext HTTP remains the default when the mode is disabled. When enabled, the endpoint serves HTTPS and can require a client certificate verified against a metrics-specific client CA.

Gateway runtime

  • Add a metrics TLS configuration table next to metrics_bind_address, with cert_path, key_path, client_ca_path, and require_client_auth.
  • Add the configuration to the file schema and CLI/file merge. Startup fails when TLS is enabled but required files are missing.
  • Replace the direct axum::serve metrics listener with an accept loop that sets TCP nodelay, applies the TLS handshake with a timeout, and serves the metrics router through hyper, following the main listener pattern.
  • Construct a dedicated TlsAcceptor for the metrics listener and use the existing reload worker. Reload logs and OCSF messages identify the listener that reloaded so gateway and metrics certificate events are distinguishable.

Trust boundary

The metrics client CA must be distinct from the gateway client CA. Sandbox pods mount the gateway client certificate; trusting the gateway client CA for metrics would allow sandbox client certificates to scrape the endpoint. The metrics listener may use the existing gateway server certificate by default, since its Service DNS SANs already cover the metrics Service identity.

Helm chart

  • Add metrics.tls.enabled with default false, plus certSecretName, clientCaSecretName, and requireClientCert values.
  • Render the metrics TLS configuration and mount referenced Secrets read-only in the gateway workload.
  • Start with bring-your-own Secrets. Chart-generated metrics CA and scraper certificates can be considered separately.
  • Offer an optional NetworkPolicy that limits metrics-port ingress to a configured namespace or pod selector as defense in depth.

This issue defines the runtime security capability and its Helm configuration. An optional, disabled-by-default ServiceMonitor may be added later, with a tlsConfig that references the scraper Secret. That work must coordinate with #2507 rather than duplicate its broader monitoring-resource scope.

Acceptance Criteria

  • The default Helm deployment retains its existing plaintext metrics behavior when metrics.tls.enabled=false.
  • When metrics.tls.enabled=true, /metrics is served over HTTPS and plaintext requests are rejected.
  • A secure configuration fails startup if required certificate, key, or client-CA files are missing.
  • A monitoring client validates the configured server certificate chain and Service DNS identity without insecureSkipVerify.
  • When client authentication is required, a client without a certificate or with a certificate not trusted by the configured metrics client CA cannot retrieve metrics.
  • A client with a certificate trusted by the configured metrics client CA retrieves metrics successfully through mTLS.
  • The metrics client CA is distinct from the gateway client CA.
  • Metrics certificate and key material is supplied through standard Kubernetes Secret references and mounted read-only, without requiring a specific issuer, monitoring controller, or Kubernetes distribution.
  • Certificate reload retains the last-known-good secure configuration if a replacement cannot be loaded, and identifies the affected listener in logs and OCSF events.
  • Runtime, Helm-rendering, and integration tests cover plaintext default behavior; secure-mode rejection of plaintext, missing, and untrusted clients; successful trusted-client scraping; and certificate reload.
  • Documentation describes the metrics TLS configuration, the separate client-CA trust boundary, certificate rotation, and a verified scrape example without verification bypasses.

Alternatives Considered

Make secure metrics the default

The currently exposed metrics are low sensitivity, and plaintext on a dedicated ClusterIP port is conventional for Prometheus scraping. Making TLS and mTLS mandatory would introduce certificate-management requirements and change the behavior of existing deployments. The secure mode should therefore be opt-in.

NetworkPolicy only

NetworkPolicy can reduce the set of workloads that reach the endpoint, but it does not encrypt traffic, verify the server identity, or authenticate the monitoring client. It remains useful as defense in depth.

HTTPS without client authentication

HTTPS protects confidentiality and verifies the server, but any workload with network access could still read the metrics. Operators that enable the secure mode need the option to require authenticated scrape access as well.

Authenticate metrics with Kubernetes ServiceAccount tokens

This would tie the gateway runtime to Kubernetes-specific token validation, audience handling, and authorization semantics. mTLS uses standard TLS primitives and works with compatible monitoring clients across deployment environments.

Terminate TLS in a sidecar or external proxy

A proxy introduces a separate certificate lifecycle, deployment path, failure mode, and security boundary. The endpoint itself should own its transport and client-authentication contract.

Expose metrics through the main gateway listener

The main listener multiplexes gateway API traffic and has a broader authentication and routing role. Retaining a dedicated metrics listener preserves a narrow monitoring boundary.

Agent Investigation

  • The Helm chart enables the dedicated metrics listener through service.metricsPort, which defaults to 9090 in values.yaml.
  • Helm renders metrics_bind_address = "0.0.0.0:<metricsPort>" in gateway-config.yaml, and the gateway Service exposes the named metrics port in service.yaml.
  • The gateway binds a raw TcpListener and serves the metrics router directly in lib.rs. metrics_router exposes GET /metrics without TLS or authentication middleware.
  • The chart contains no ServiceMonitor, PodMonitor, tlsConfig, or insecureSkipVerify configuration.
  • OpenShell already provides reusable TLS, mTLS client verification, certificate reload, and last-known-good configuration behavior through TlsAcceptor.
  • A deployment of the current chart confirmed the runtime behavior described above: the Service exposes the metrics port, GET /metrics returns metrics over plaintext HTTP, an invalid Authorization header does not restrict access, and an HTTPS connection cannot be established.

Related work: #2507 and #3179. #2507 covers broader observability and may later build on an optional ServiceMonitor for this secure endpoint; #3179 covers listener-port validation. Neither implements the secure metrics mode described here.

Activity

  1. changed the title [-]feat(metrics): secure Prometheus metrics endpoint with mTLS[/-] [+]feat(metrics): secure gateway metrics endpoint with mTLS[/+] on Sep 24, 2026
  2. krishicks commented on Sep 24, 2026

    @krishicks
    Collaborator

    I think this is a reasonable addition.

    🤖 implementation plan follows:

    The current metrics content is low-sensitivity: gRPC method and status counts, latencies, normalized HTTP paths, database readiness, and interceptor names and binding IDs. It contains no secrets, identities, or sandbox names. Sandbox workloads also cannot reach the metrics port unless their policy allows it. Plaintext on a dedicated ClusterIP port is the standard Prometheus convention, so this should be an opt-in feature, with plaintext remaining the default. The existing TLS machinery covers most of the work.

    Gateway runtime

    • Config: add a metrics TLS table (cert_path, key_path, client_ca_path, require_client_auth) next to metrics_bind_address in openshell-core/src/config.rs. It also goes in the file schema (openshell-server/src/config_file.rs) and the CLI/file merge (openshell-server/src/cli.rs). Startup fails if TLS is enabled but the files are missing.
    • Listener: lib.rs currently passes a plain TcpListener to axum::serve. Replace that with an accept loop: accept, set TCP nodelay, run the TLS handshake with a timeout, then serve the metrics router with hyper. This follows the main listener's pattern.
    • Reload: TlsAcceptor::from_files works as-is with None for the external-cert arguments. It already provides mTLS client verification, hot reload, and last-known-good retention. The metrics listener constructs its own instance and calls spawn_reload_worker on shutdown. Reload log and OCSF messages in tls.rs should identify which listener reloaded so operators can tell the gateway and metrics certificates apart.

    Trust boundary

    The metrics client CA must be separate from the gateway client CA. Sandbox pods mount the gateway client certificate (server.tls.clientTlsSecretName). If the metrics listener trusted that CA, any sandbox client certificate could scrape metrics. The server certificate can default to the existing gateway server certificate, since its SANs already cover the Service DNS name.

    Helm chart

    • Values: metrics.tls.enabled (default false), certSecretName, clientCaSecretName, and requireClientCert.
    • Mounts: read-only Secret volume mounts in _gateway-workload.tpl and rendering in gateway-config.yaml.
    • PKI: start with bring-your-own Secrets. Chart-generated metrics client CA and scraper certs (via cert-manager-pki.yaml, with a matching certgen.yaml path) can follow separately.
    • NetworkPolicy: an optional policy that restricts ingress on the metrics port to a configured namespace or pod selector. The existing networkpolicy.yaml only covers sandbox SSH.
    • Tests: Helm unit tests in tests/gateway_config_test.yaml.

    Tests

    • Server unit tests: plaintext request rejected, missing client cert rejected, untrusted client cert rejected, trusted client cert returns metrics, and the served certificate changes after reload. tls.rs already has cert-generation helpers.
    • E2E: one kind/Helm scenario that scrapes with a client certificate.

    Docs

    • docs/reference/gateway-config.mdx for the new TOML table.
    • The Helm chart README.
    • A short note in the relevant architecture/ doc on the metrics listener trust boundary.
    • skills/debug-openshell-cluster/SKILL.md for the chart changes.
    • An example Prometheus Operator tlsConfig that references the scraper Secret. Secrets referenced from a ServiceMonitor/PodMonitor must be in the monitor's namespace.

    Optional extra

    The chart could ship an optional ServiceMonitor for the metrics endpoint, disabled by default. When metrics TLS is enabled, it would come preconfigured with a tlsConfig that references the scraper Secret. The broader monitoring-resource work in #2507 should build on it rather than duplicate it.

    Size

    The runtime change is small: a few hundred lines plus tests. The Helm and PKI work is the larger share. A first PR could be limited to the runtime change, bring-your-own Secrets, and the optional NetworkPolicy.

  3. changed the title [-]feat(metrics): secure gateway metrics endpoint with mTLS[/-] [+]feat(metrics): opt-in secure gateway metrics endpoint[/+] on Sep 25, 2026
  4. gmenher commented on Sep 25, 2026

    @gmenher
    ContributorAuthor

    Thanks @krishicks, makes sense. Agreed that the current metrics surface is low sensitivity and that plaintext on a dedicated port is a conventional default prometheus setup, so the secure endpoint should be opt-in rather than changing existing default behavior.
    I've updated thIS issue to reflect the new direction and the implementation guidance, including the separate metrics client-CA trust boundary, bring-your-own Secrets as the initial chart scope, optional NetworkPolicy, and coordination with #2507 for any future ServiceMonitor work. Thanks for the feedback and the clarification.

  5. kvnloo commented on Sep 27, 2026

    @kvnloo

    The revised opt-in direction looks good.

    One invariant I would make explicit in tests is that enabling metrics mTLS must not accidentally trust the normal gateway client CA.

    A useful trust matrix:

    • plaintext/default mode -> existing behavior unchanged;
    • TLS without client auth -> valid server TLS, anonymous scrape allowed;
    • mTLS + metrics client CA -> metrics client cert accepted;
    • sandbox/gateway client cert signed by the normal gateway client CA -> rejected;
    • invalid replacement certificate on reload -> existing last-known-good listener remains usable.

    That last pair seems important because reusing the normal gateway CA would silently allow sandbox identities into the metrics endpoint, while a failed reload should not turn observability configuration into a gateway availability event.

    I would also keep chart-generated certificate lifecycle out of the first PR and use BYO Secrets as proposed. That keeps runtime TLS semantics independently reviewable.

  6. gmenher commented on Sep 28, 2026

    @gmenher
    ContributorAuthor

    @krishicks , I would be happy to start working on this. If you think it makes sense, I can start the first small limited PR for the runtime TLS/mTLS path and bring-your-own Secrets while leaving the rest separate for now.

  7. added
    state:acceptedA maintainer decided OpenShell should pursue this issue
    and removed
    state:triage-neededOpened without agent diagnostics and needs triage
    on Sep 28, 2026
  8. gmenher commented on Oct 6, 2026

    @gmenher
    ContributorAuthor

    Hi @krishicks, I’ve opened and rebased #3877 onto the latest main and resolved the conflicts.

    This is the first implementation PR for this issue. It adds opt-in TLS/mTLS for the dedicated metrics listener, using operator-provided Secrets in Helm, while keeping plaintext metrics as the default. It also keeps the metrics client CA separate from the gateway client CA so sandbox-held credentials cannot scrape metrics.

    I'd greatly appreciate a review on it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    state:acceptedA maintainer decided OpenShell should pursue this issue

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions