> ## Documentation Index
> Fetch the complete documentation index at: https://nestcrate.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Add metrics and tracing to nestrs apps

> Wire up structured request logs, Prometheus RED metrics, and optional OpenTelemetry trace/metric/log export in a few builder calls on NestApplication.

nestrs builds its observability story around four composable layers: a global `tracing` subscriber for structured logging, a request-tracing middleware that emits per-request spans and log lines, an optional Prometheus scrape endpoint, and — behind the `otel` feature flag — OpenTelemetry export for **traces, metrics, and logs** that can be layered on without changing the rest of the pipeline.

## Installing the tracing subscriber

Call `configure_tracing` once, before `listen`, so all log output and request spans share the same pipeline. The log level is read from `NESTRS_LOG` first, then `RUST_LOG`, then the `default_directive` on `TracingConfig` (default `"info"`):

```rust theme={null}
use nestrs::prelude::*;

#[module]
struct AppModule;

#[tokio::main]
async fn main() {
    let tracing = TracingConfig::builder()
        .format(TracingFormat::Json)           // Pretty for local dev
        .default_directive("info,nestrs=debug");

    NestFactory::create::<AppModule>()
        .configure_tracing(tracing)
        .listen(3000)
        .await;
}
```

`TracingFormat::Pretty` is the default and produces human-readable multi-line output. Switch to `TracingFormat::Json` in production for log aggregation platforms.

## Request tracing middleware

`use_request_tracing` adds a middleware that:

* Records a completion log line with `method`, `path`, `status`, `duration_ms`, and `request_id` (when `use_request_id()` is also enabled).
* Creates a `tracing` span named `http.server.request` for each request, with fields `http.request.method` and `http.route`.

Skip high-volume infrastructure paths (metrics, health) to avoid flooding request logs:

```rust theme={null}
NestFactory::create::<AppModule>()
    .configure_tracing(TracingConfig::builder())
    .use_request_id()
    .use_request_tracing(
        RequestTracingOptions::builder()
            .skip_paths(["/metrics", "/health"])
    )
    .enable_metrics("/metrics")
    .enable_health_check("/health")
    .listen(3000)
    .await;
```

<Note>
  `http.route` is set to the concrete request path at this middleware layer. Axum's route template (e.g. `/users/:id`) is not available here. For OTLP dashboards, treat the literal path as the closest stable route identifier unless you add a custom layer that sets a template field.
</Note>

## Prometheus metrics

`enable_metrics` registers a Prometheus scrape handler at the path you provide (default `/metrics`). It tracks:

* `http_request_duration_seconds` — histogram with standard buckets
* `http_requests_total{method, status}` — counter
* `http_requests_in_flight` — in-flight gauge

```rust theme={null}
NestFactory::create::<AppModule>()
    .use_request_tracing(RequestTracingOptions::builder().skip_paths(["/metrics"]))
    .enable_metrics("/metrics")
    .listen(3000)
    .await;
```

The `/metrics` path is mounted at the server root — it is not affected by `set_global_prefix` or `enable_uri_versioning`.

## Health and readiness checks

Use `enable_health_check` for a simple liveness probe that always returns `200`:

```rust theme={null}
.enable_health_check("/health")
```

Use `enable_readiness_check` when you want to gate traffic on the health of dependencies. Implement `HealthIndicator` for each dependency and pass the indicators at startup:

```rust theme={null}
use nestrs::prelude::*;
use std::sync::Arc;

pub struct DatabaseHealth {
    pool: Arc<sqlx::PgPool>,
}

#[async_trait]
impl HealthIndicator for DatabaseHealth {
    fn name(&self) -> &'static str { "database" }

    async fn check(&self) -> HealthStatus {
        match sqlx::query("SELECT 1").execute(self.pool.as_ref()).await {
            Ok(_) => HealthStatus::Up,
            Err(e) => HealthStatus::down(e.to_string()),
        }
    }
}

// In main:
let db_health = Arc::new(DatabaseHealth { pool: pool.clone() });

NestFactory::create::<AppModule>()
    .enable_readiness_check("/ready", [db_health as Arc<dyn HealthIndicator>])
    .listen(3000)
    .await;
```

When any indicator returns `HealthStatus::Down`, the readiness endpoint returns `503` with a Terminus-style JSON summary containing `status`, `info`, `error`, and `details` keys.

## Probe decorators (`#[liveness]`, `#[readiness]`, `#[startup]`)

Beyond the always-OK liveness route, nestrs can mirror a *real handler* as a probe. Decorating a route handler with one of the probe decorators mounts a fixed endpoint under `/__nestrs/health/*` that calls the handler on each scrape (or once, for startup) and reports up/down from its status code:

| Decorator | Mirrored endpoint | Semantics |
| - | - | - |
| `#[liveness]` | `GET /__nestrs/health/live` | 2xx from the handler ⇒ up, otherwise down. |
| `#[readiness]` | `GET /__nestrs/health/ready` | 2xx ⇒ up, otherwise down. |
| `#[startup]` | `GET /__nestrs/health/startup` | Evaluated **once per process** and cached; later probes return the first result. |

```rust theme={null}
use nestrs::prelude::*;

#[routes(state = AppState)]
impl HealthController {
    /// Real dependency checks — the probe endpoint runs this handler.
    #[get("/health/deep")]
    #[readiness]
    pub async fn deep(State(s): State<AppState>) -> impl IntoResponse {
        match s.check_deps().await {
            Ok(_) => (StatusCode::OK, Json(json!({"deps": "ok"}))),
            Err(_) => (StatusCode::SERVICE_UNAVAILABLE, Json(json!({"deps": "down"}))),
        }
    }
}
```

The mirrored endpoints are mounted at the **server root** (not under `set_global_prefix` or URI versioning) so orchestrators can probe them without prefixes. Probe results are cached for 5 seconds per endpoint to prevent probe-storm self-DoS; a panicking handler reports down with a generic message (raw errors go to `tracing` only).

<Note>
  If your own routes already occupy <code>/\_\_nestrs/health/live</code> or <code>/\_\_nestrs/health/ready</code>, your routes win — the framework skips its mirrored endpoint. The builder-level <code>enable\_health\_check("/health")</code> and <code>enable\_readiness\_check("/ready", \[...])</code> paths above are unaffected and remain the simplest options for most apps.
</Note>

Use decorator-based probes when liveness should depend on something the app actually does (a warmup pass, a self-check handler), and the builder calls when a constant 200 plus indicator checks is enough. Kubernetes wiring: `livenessProbe` → `/__nestrs/health/live` (or your `enable_health_check` path), `readinessProbe` → `/__nestrs/health/ready` (or your readiness path), `startupProbe` → `/__nestrs/health/startup`.

## OpenTelemetry (OTLP)

Enable the `otel` feature to export to any OTLP-compatible collector (Jaeger, Tempo, Honeycomb, SigNoz, etc.). **Spans** export by default; **metrics** and **logs** are per-signal opt-ins on the config:

```toml theme={null}
[dependencies]
nestrs = { version = "1.6.0", features = ["otel"] }
```

Replace `configure_tracing` with `configure_tracing_opentelemetry` and supply an `OpenTelemetryConfig`:

```rust theme={null}
use nestrs::prelude::*;

#[module]
struct AppModule;

#[tokio::main]
async fn main() {
    let tracing = TracingConfig::builder().format(TracingFormat::Json);
    let otel = OpenTelemetryConfig::new("my-service")
        .endpoint("http://localhost:4317")
        .sample_ratio(1.0)
        .metrics()   // ← also push metrics over OTLP
        .logs();     // ← also push logs over OTLP

    NestFactory::create::<AppModule>()
        .configure_tracing_opentelemetry(tracing, otel)
        .use_request_id()
        .use_request_tracing(RequestTracingOptions::builder().skip_paths(["/metrics"]))
        .enable_metrics("/metrics")
        .listen(3000)
        .await;
}
```

When `endpoint` is not set, nestrs falls back to the `OTEL_EXPORTER_OTLP_ENDPOINT` environment variable, then to `http://localhost:4317`.

### Dual metrics export: Prometheus pull + OTLP push

There is exactly **one** metrics recording surface in the process — the `metrics` facade — and nestrs owns it with a fan-out recorder. Every `metrics::counter!` / `gauge!` / `histogram!` call is forwarded to **every** enabled backend:

* `enable_metrics("/metrics")` registers the **Prometheus** backend (pull, scraped at the path you choose).
* `OpenTelemetryConfig::metrics()` registers the **OTLP** backend (push, on the collector's schedule).

They compose freely and in either order. An app with both enabled gets identical series in Prometheus and in the collector — the framework's RED metrics (`http_requests_total`, `http_request_duration_seconds`, `http_requests_in_flight`) *and* any instruments you define yourself. Use this to bridge a migration (Prometheus today, OTLP tomorrow) or to serve both a local scrape and a central pipeline.

<Note>
  Units are translated to OpenTelemetry's UCUM-style conventions when pushed: facade <code>Unit::Seconds</code> exports as <code>s</code>, <code>Unit::Count</code> as <code>1</code>. Facade labels become OTel attributes 1:1 (<code>counter!("rps", "route" => "/x")</code> exports with attribute <code>route="/x"</code>).
</Note>

### Custom instruments

Record your own metrics through the same surface — the `metrics` facade is re-exported at `nestrs::metrics`, so no extra dependency is needed:

```rust theme={null}
// Anywhere in request handlers, services, or background jobs:
nestrs::metrics::counter!("orders_created_total", "kind" => "subscription").increment(1);
nestrs::metrics::histogram!("payment_latency_seconds").record(0.42);
nestrs::metrics::gauge!("queue_depth").set(17.0);
```

Instruments declared before first use with `nestrs::metrics::describe_counter!(...)` etc. carry their unit and description into both backends.

### Logs over OTLP

`OpenTelemetryConfig::logs()` bridges the `tracing` facade into the OTLP log pipeline: every `tracing::info!` / `warn!` / `error!` event becomes an OpenTelemetry log record on the same collector as your traces. Events emitted inside an active span are correlated with its trace and span IDs, so a trace in Jaeger and its surrounding log lines in Loki/Tempo line up automatically. The local `tracing` output (pretty or JSON) is unchanged — OTLP is an additional destination, not a replacement.

### Runtime requirements and shutdown

* **Install from async context.** The OTLP exporters are lazy — a missing collector never fails startup — but their transport is spawned onto the current Tokio reactor at construction. Call `configure_tracing_opentelemetry` from `#[tokio::main]` (before `listen`), exactly as in the examples above.
* **Metric export cadence** follows the collector convention: every 60s by default, overridable with `OTEL_METRIC_EXPORT_INTERVAL` (milliseconds).
* **Shutdown flushes.** `NestApplication::listen*` methods flush and stop the OTLP pipelines on graceful shutdown, so the last metrics window and log batch reach the collector. Calling the install functions directly (`nestrs::otel::install_otlp_meter` / `install_otlp_logger`) pairs with `nestrs::otel::shutdown_meter_provider()` / `shutdown_logger_provider()` for custom lifecycles.

<Tip>
  Use `try_init_tracing` or `try_init_tracing_opentelemetry` directly if you need to install the tracing subscriber outside of the `NestApplication` builder chain (for example in test harnesses or CLI tools).
</Tip>

## Environment variables reference

| Variable | Role |
| - | - |
| `NESTRS_LOG` | Primary log filter directive. Overrides `RUST_LOG` and `TracingConfig::default_directive`. |
| `RUST_LOG` | Fallback filter when `NESTRS_LOG` is unset (standard `tracing-subscriber` semantics). |
| `OTEL_EXPORTER_OTLP_ENDPOINT` | OTLP collector address when not set via `OpenTelemetryConfig::endpoint`. |
| `OTEL_METRIC_EXPORT_INTERVAL` | OTLP metric push cadence in milliseconds (default 60000). |

## Troubleshooting

| Symptom | What to check |
| - | - |
| No log output | `configure_tracing` must run before `listen`. Check `NESTRS_LOG` / `RUST_LOG` and confirm nothing else installs a conflicting global subscriber. |
| `/metrics` floods access logs | Add `/metrics` to `RequestTracingOptions::skip_paths`. |
| Spans missing in Jaeger or Tempo | Confirm the `otel` feature is enabled, the endpoint URL is reachable, and the sampling ratio is > 0. Check that the collector receives traffic on the expected gRPC or HTTP port. |
| OTel metrics/logs missing in the collector | Confirm `.metrics()` / `.logs()` are set on the `OpenTelemetryConfig` (spans export without them; metrics and logs do not). Metric pushes are periodic — wait one interval (`OTEL_METRIC_EXPORT_INTERVAL`) before concluding failure. |
| `there is no reactor running` panic at startup | OTLP pipelines must be constructed from async context — call `configure_tracing_opentelemetry` from inside `#[tokio::main]`, not from a plain `fn main` before building the runtime. |
| High cardinality in `http.route` | Expected — path is the literal URI at this layer. Add a custom layer or a business metric if you need route-template labels in OTLP dashboards. |

## Local dev vs production

<Tabs>
  <Tab title="Local development">
    ```rust theme={null}
    TracingConfig::builder()
        .format(TracingFormat::Pretty)
        .default_directive("debug,nestrs=trace")
    ```
  </Tab>

  <Tab title="Production">
    ```rust theme={null}
    let tracing = TracingConfig::builder()
        .format(TracingFormat::Json)
        .default_directive("info");

    let otel = OpenTelemetryConfig::new("my-service")
        .sample_ratio(0.1);  // sample 10% in high-volume prod
    ```
  </Tab>
</Tabs>
