Repository navigation
feat(metrics): opt-in secure gateway metrics endpoint #3664
Description
Activity
- addedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 24, 2026 - changed the title
[-]feat(metrics): secure Prometheus metrics endpoint with mTLS[/-][+]feat(metrics): secure gateway metrics endpoint with mTLS[/+]on Sep 24, 2026 I think this is a reasonable addition.
🤖 implementation plan follows:
The current metrics content is low-sensitivity: gRPC method and status counts, latencies, normalized HTTP paths, database readiness, and interceptor names and binding IDs. It contains no secrets, identities, or sandbox names. Sandbox workloads also cannot reach the metrics port unless their policy allows it. Plaintext on a dedicated ClusterIP port is the standard Prometheus convention, so this should be an opt-in feature, with plaintext remaining the default. The existing TLS machinery covers most of the work.
Gateway runtime
- Config: add a metrics TLS table (
cert_path,key_path,client_ca_path,require_client_auth) next tometrics_bind_addressinopenshell-core/src/config.rs. It also goes in the file schema (openshell-server/src/config_file.rs) and the CLI/file merge (openshell-server/src/cli.rs). Startup fails if TLS is enabled but the files are missing. - Listener:
lib.rscurrently passes a plainTcpListenertoaxum::serve. Replace that with an accept loop: accept, set TCP nodelay, run the TLS handshake with a timeout, then serve the metrics router with hyper. This follows the main listener's pattern. - Reload:
TlsAcceptor::from_filesworks as-is withNonefor the external-cert arguments. It already provides mTLS client verification, hot reload, and last-known-good retention. The metrics listener constructs its own instance and callsspawn_reload_workeron shutdown. Reload log and OCSF messages intls.rsshould identify which listener reloaded so operators can tell the gateway and metrics certificates apart.
Trust boundary
The metrics client CA must be separate from the gateway client CA. Sandbox pods mount the gateway client certificate (
server.tls.clientTlsSecretName). If the metrics listener trusted that CA, any sandbox client certificate could scrape metrics. The server certificate can default to the existing gateway server certificate, since its SANs already cover the Service DNS name.Helm chart
- Values:
metrics.tls.enabled(defaultfalse),certSecretName,clientCaSecretName, andrequireClientCert. - Mounts: read-only Secret volume mounts in
_gateway-workload.tpland rendering ingateway-config.yaml. - PKI: start with bring-your-own Secrets. Chart-generated metrics client CA and scraper certs (via
cert-manager-pki.yaml, with a matchingcertgen.yamlpath) can follow separately. - NetworkPolicy: an optional policy that restricts ingress on the metrics port to a configured namespace or pod selector. The existing
networkpolicy.yamlonly covers sandbox SSH. - Tests: Helm unit tests in
tests/gateway_config_test.yaml.
Tests
- Server unit tests: plaintext request rejected, missing client cert rejected, untrusted client cert rejected, trusted client cert returns metrics, and the served certificate changes after reload.
tls.rsalready has cert-generation helpers. - E2E: one kind/Helm scenario that scrapes with a client certificate.
Docs
docs/reference/gateway-config.mdxfor the new TOML table.- The Helm chart README.
- A short note in the relevant
architecture/doc on the metrics listener trust boundary. skills/debug-openshell-cluster/SKILL.mdfor the chart changes.- An example Prometheus Operator
tlsConfigthat references the scraper Secret. Secrets referenced from aServiceMonitor/PodMonitormust be in the monitor's namespace.
Optional extra
The chart could ship an optional
ServiceMonitorfor the metrics endpoint, disabled by default. When metrics TLS is enabled, it would come preconfigured with atlsConfigthat references the scraper Secret. The broader monitoring-resource work in #2507 should build on it rather than duplicate it.Size
The runtime change is small: a few hundred lines plus tests. The Helm and PKI work is the larger share. A first PR could be limited to the runtime change, bring-your-own Secrets, and the optional NetworkPolicy.
- Config: add a metrics TLS table (
- changed the title
[-]feat(metrics): secure gateway metrics endpoint with mTLS[/-][+]feat(metrics): opt-in secure gateway metrics endpoint[/+]on Sep 25, 2026 Thanks @krishicks, makes sense. Agreed that the current metrics surface is low sensitivity and that plaintext on a dedicated port is a conventional default prometheus setup, so the secure endpoint should be opt-in rather than changing existing default behavior.
I've updated thIS issue to reflect the new direction and the implementation guidance, including the separate metrics client-CA trust boundary, bring-your-own Secrets as the initial chart scope, optional NetworkPolicy, and coordination with #2507 for any future ServiceMonitor work. Thanks for the feedback and the clarification.The revised opt-in direction looks good.
One invariant I would make explicit in tests is that enabling metrics mTLS must not accidentally trust the normal gateway client CA.
A useful trust matrix:
- plaintext/default mode -> existing behavior unchanged;
- TLS without client auth -> valid server TLS, anonymous scrape allowed;
- mTLS + metrics client CA -> metrics client cert accepted;
- sandbox/gateway client cert signed by the normal gateway client CA -> rejected;
- invalid replacement certificate on reload -> existing last-known-good listener remains usable.
That last pair seems important because reusing the normal gateway CA would silently allow sandbox identities into the metrics endpoint, while a failed reload should not turn observability configuration into a gateway availability event.
I would also keep chart-generated certificate lifecycle out of the first PR and use BYO Secrets as proposed. That keeps runtime TLS semantics independently reviewable.
Reacted by Gaizka Menendez@krishicks , I would be happy to start working on this. If you think it makes sense, I can start the first small limited PR for the runtime TLS/mTLS path and bring-your-own Secrets while leaving the rest separate for now.
Reacted by krishicks- addedstate:acceptedA maintainer decided OpenShell should pursue this issueA maintainer decided OpenShell should pursue this issueand removedstate:triage-neededOpened without agent diagnostics and needs triageOpened without agent diagnostics and needs triage
on Sep 28, 2026 Hi @krishicks, I’ve opened and rebased #3877 onto the latest main and resolved the conflicts.
This is the first implementation PR for this issue. It adds opt-in TLS/mTLS for the dedicated metrics listener, using operator-provided Secrets in Helm, while keeping plaintext metrics as the default. It also keeps the metrics client CA separate from the gateway client CA so sandbox-held credentials cannot scrape metrics.
I'd greatly appreciate a review on it.
User Story
As an OpenShell platform operator, I want to opt in to a gateway metrics endpoint that uses verified HTTPS and an authenticated monitoring client, so that I can collect operational metrics under a defined transport and trust contract without disabling certificate verification.
Problem Statement
The Helm chart enables a dedicated gateway metrics listener on port
9090by default and exposes it through the OpenShell Service as the namedmetricsport.The gateway serves
GET /metricsdirectly over a raw TCP listener. The endpoint does not perform a TLS handshake and does not authenticate or authorize the caller. A workload that can reach the Service can read the metrics over plaintext HTTP.There is currently no metrics TLS configuration, client-authentication configuration, or secure scrape configuration in the chart. Network reachability can be constrained by deployment policy, but the endpoint itself does not provide transport confidentiality, server identity verification, or caller authentication.
Impact / Why This Matters
The currently emitted metrics are low-sensitivity operational signals: gRPC method and status counts, request latencies, normalized HTTP paths, database readiness, and gateway-interceptor names and binding identifiers. They do not contain secrets, identities, or sandbox names. Plaintext scraping on a dedicated ClusterIP port is therefore a valid default Prometheus workflow.
Some operators nevertheless require encrypted transport, verified server identity, and authenticated scrape access for every observability endpoint. Today they can disable the metrics endpoint or rely on deployment-specific network policy, but neither option provides a supported secure scrape contract. An opt-in TLS and mTLS mode gives those operators a portable solution while preserving the default behavior for existing deployments.
Proposed Design
OpenShell should add an optional secure metrics mode. Plaintext HTTP remains the default when the mode is disabled. When enabled, the endpoint serves HTTPS and can require a client certificate verified against a metrics-specific client CA.
Gateway runtime
metrics_bind_address, withcert_path,key_path,client_ca_path, andrequire_client_auth.axum::servemetrics listener with an accept loop that sets TCP nodelay, applies the TLS handshake with a timeout, and serves the metrics router through hyper, following the main listener pattern.TlsAcceptorfor the metrics listener and use the existing reload worker. Reload logs and OCSF messages identify the listener that reloaded so gateway and metrics certificate events are distinguishable.Trust boundary
The metrics client CA must be distinct from the gateway client CA. Sandbox pods mount the gateway client certificate; trusting the gateway client CA for metrics would allow sandbox client certificates to scrape the endpoint. The metrics listener may use the existing gateway server certificate by default, since its Service DNS SANs already cover the metrics Service identity.
Helm chart
metrics.tls.enabledwith defaultfalse, pluscertSecretName,clientCaSecretName, andrequireClientCertvalues.This issue defines the runtime security capability and its Helm configuration. An optional, disabled-by-default ServiceMonitor may be added later, with a
tlsConfigthat references the scraper Secret. That work must coordinate with #2507 rather than duplicate its broader monitoring-resource scope.Acceptance Criteria
metrics.tls.enabled=false.metrics.tls.enabled=true,/metricsis served over HTTPS and plaintext requests are rejected.insecureSkipVerify.Alternatives Considered
Make secure metrics the default
The currently exposed metrics are low sensitivity, and plaintext on a dedicated ClusterIP port is conventional for Prometheus scraping. Making TLS and mTLS mandatory would introduce certificate-management requirements and change the behavior of existing deployments. The secure mode should therefore be opt-in.
NetworkPolicy only
NetworkPolicy can reduce the set of workloads that reach the endpoint, but it does not encrypt traffic, verify the server identity, or authenticate the monitoring client. It remains useful as defense in depth.
HTTPS without client authentication
HTTPS protects confidentiality and verifies the server, but any workload with network access could still read the metrics. Operators that enable the secure mode need the option to require authenticated scrape access as well.
Authenticate metrics with Kubernetes ServiceAccount tokens
This would tie the gateway runtime to Kubernetes-specific token validation, audience handling, and authorization semantics. mTLS uses standard TLS primitives and works with compatible monitoring clients across deployment environments.
Terminate TLS in a sidecar or external proxy
A proxy introduces a separate certificate lifecycle, deployment path, failure mode, and security boundary. The endpoint itself should own its transport and client-authentication contract.
Expose metrics through the main gateway listener
The main listener multiplexes gateway API traffic and has a broader authentication and routing role. Retaining a dedicated metrics listener preserves a narrow monitoring boundary.
Agent Investigation
service.metricsPort, which defaults to9090invalues.yaml.metrics_bind_address = "0.0.0.0:<metricsPort>"ingateway-config.yaml, and the gateway Service exposes the namedmetricsport inservice.yaml.TcpListenerand serves the metrics router directly inlib.rs.metrics_routerexposesGET /metricswithout TLS or authentication middleware.ServiceMonitor,PodMonitor,tlsConfig, orinsecureSkipVerifyconfiguration.TlsAcceptor.GET /metricsreturns metrics over plaintext HTTP, an invalidAuthorizationheader does not restrict access, and an HTTPS connection cannot be established.Related work: #2507 and #3179. #2507 covers broader observability and may later build on an optional ServiceMonitor for this secure endpoint; #3179 covers listener-port validation. Neither implements the secure metrics mode described here.