Writing

Building an OpenTelemetry Pipeline with Grafana Cloud

Building an OpenTelemetry pipeline, from the first trace to metrics, logs, and remote Collector management

How to Build an OpenTelemetry Pipeline with Grafana Cloud

An API request may be handled by several functions, services, and databases before a response is returned. If 900 ms is reported in a completion log, the slow operation is not identified. With tracing, the time spent in each operation is recorded.

A trace is created in Go and sent to an OpenTelemetry Collector. Metrics and logs are then added to the same receiver. The three signals are routed to Grafana Cloud, and Collector configuration is managed through OpAMP and Grafana Fleet Management.

Contents

Starting with one span

A span is used to record one operation and its duration. Attributes, events, and an error status can also be attached. Related spans are connected in a trace. The following trace could be produced by one request:

HandleRequest                  900 ms
├── ValidateInput               40 ms
├── CallDependency             620 ms
└── SaveResult                 180 ms

To connect these operations in one trace, a span is started, its context is passed to the next operation, and each span is ended when the work is finished:

ctx, span := tracer.Start(ctx, "HandleRequest")
defer span.End()

ctx, dependencySpan := tracer.Start(ctx, "CallDependency")
err := callDependency(ctx)
if err != nil {
	dependencySpan.RecordError(err)
}
dependencySpan.End()

The context returned by HandleRequest is passed to the child span. The trace relationship is carried in that context. Across service boundaries, the same trace context is carried in request headers by instrumentation or propagators on both sides.

The spans are next exported from the Go process.

Sending spans with OTLP

APIs and SDKs for creating telemetry are provided by OpenTelemetry, often called OTel. Telemetry is sent with OTLP. Spans can be sent by the Go SDK over OTLP/gRPC to a local receiver on port 4317.

The Go tracer provider is connected to the receiver through an exporter. The following excerpt is placed in application startup, such as main. The application and receiver are assumed to run on the same machine:

ctx := context.Background()

exporter, err := otlptracegrpc.New(
	ctx,
	otlptracegrpc.WithEndpoint("127.0.0.1:4317"),
	otlptracegrpc.WithInsecure(),
)
if err != nil {
	log.Fatal(err)
}

provider := sdktrace.NewTracerProvider(
	sdktrace.WithBatcher(exporter),
	sdktrace.WithResource(
		resource.NewSchemaless(
			attribute.String("service.name", "example-service"),
		),
	),
)
defer provider.Shutdown(context.Background())

tracer := provider.Tracer("example-service")

// Application work is started after the tracer is configured.

Finished spans are grouped before export by WithBatcher. The producing service is identified by service.name. Spans still queued when the process ends are sent during shutdown. A local connection to the Collector is used in this example.

Spans are now sent by the application. A receiver is needed at 127.0.0.1:4317.

Receiving and inspecting spans with a Collector

Telemetry is received, processed, and exported by the OpenTelemetry Collector, a separate process. Each Collector pipeline is started with a receiver and ended with an exporter. Processors can be placed between them.

Before data are forwarded to Grafana Cloud, the arrival of spans at the Collector can be checked with a debug exporter. The following trace pipeline can be used for a local test:

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 127.0.0.1:4317

exporters:
  debug:
    verbosity: detailed

service:
  pipelines:
    traces:
      receivers: [otlp]
      exporters: [debug]

The application’s spans are accepted by the otlp receiver. Information about those spans is written to the Collector’s output by the debug exporter. This output is used to check the first hop:

Application ──OTLP/gRPC──▶ Collector ──▶ debug output

Once the first span has been received, metrics and logs can be added to the same OTLP receiver.

Adding metrics when one trace is not enough

The time spent by one request is shown by a trace. To see changes across the service, metrics are recorded over time. Completed requests can be counted with a counter, and their durations can be recorded with a histogram:

request count:       120 → 150 → 190
request duration:    80 ms, 95 ms, 900 ms, ...

A duration range or percentile can be calculated from the histogram. The particular 900 ms request can still be inspected through its trace. A metric exporter and meter provider are added beside the trace exporter, and the same local Collector endpoint is used:

metricExporter, err := otlpmetricgrpc.New(
	ctx,
	otlpmetricgrpc.WithEndpoint("127.0.0.1:4317"),
	otlpmetricgrpc.WithInsecure(),
)
if err != nil {
	log.Fatal(err)
}

meterProvider := sdkmetric.NewMeterProvider(
	sdkmetric.WithReader(sdkmetric.NewPeriodicReader(metricExporter)),
	sdkmetric.WithResource(
		resource.NewSchemaless(attribute.String("service.name", "example-service")),
	),
)
defer meterProvider.Shutdown(context.Background())

meter := meterProvider.Meter("example-service")
requests, err := meter.Int64Counter("requests")
if err != nil {
	log.Fatal(err)
}

The counter is created during startup. In a request handler, its value is increased using the request context:

requests.Add(ctx, 1)

The metric is exported through OTLP by the periodic reader.

Adding logs for event details

An increase in errors can be shown by a metric, and one failed request can be located with a trace. The details of an event are recorded in a log:

time=2026-09-27T10:05:00Z level=error service=example-service message="dependency timed out" trace_id=...

A log is linked to a trace when the active trace context is included in the record. A Go slog logger can be connected to an OpenTelemetry log provider, and the records can be sent to the same Collector:

logExporter, err := otlploggrpc.New(
	ctx,
	otlploggrpc.WithEndpoint("127.0.0.1:4317"),
	otlploggrpc.WithInsecure(),
)
if err != nil {
	log.Fatal(err)
}

logProvider := sdklog.NewLoggerProvider(
	sdklog.WithProcessor(sdklog.NewBatchProcessor(logExporter)),
	sdklog.WithResource(
		resource.NewSchemaless(attribute.String("service.name", "example-service")),
	),
)
defer logProvider.Shutdown(context.Background())

logger := otelslog.NewLogger(
	"example-service",
	otelslog.WithLoggerProvider(logProvider),
)

In a request handler, an error is recorded with the context returned by tracer.Start:

logger.ErrorContext(ctx, "dependency timed out")

The log can then be related to the active trace. The same service.name is attached to all three signals. Logs produced by other sources can be received through a separate Collector receiver.

Sending all three signals to Grafana Cloud

Traces, metrics, and logs are received on the Collector’s OTLP/gRPC endpoint. They are sent to Grafana Cloud through one otlphttp/grafana_cloud exporter. A separate Collector pipeline is configured for each signal, and the same Grafana Cloud OTLP endpoint is used by all three. The endpoint and credentials are obtained from the Grafana Cloud stack.

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 127.0.0.1:4317

processors:
  batch:

extensions:
  basicauth/grafana_cloud:
    client_auth:
      username: ${env:GRAFANA_CLOUD_INSTANCE_ID}
      password: ${env:GRAFANA_CLOUD_TOKEN}

exporters:
  otlphttp/grafana_cloud:
    endpoint: ${env:GRAFANA_CLOUD_OTLP_ENDPOINT}
    auth:
      authenticator: basicauth/grafana_cloud

service:
  extensions: [basicauth/grafana_cloud]
  pipelines:
    traces:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlphttp/grafana_cloud]
    metrics:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlphttp/grafana_cloud]
    logs:
      receivers: [otlp]
      processors: [batch]
      exporters: [otlphttp/grafana_cloud]

The OTLP URL, instance ID, and access token are supplied through three environment variables. A path for each signal is appended to the /otlp URL by the otlphttp exporter. The Collector’s basic-auth extension is used to send the credentials.

Application ──OTLP──▶ Collector
                         ├── traces ──▶ Grafana Cloud OTLP endpoint
                         ├── metrics ─▶ Grafana Cloud OTLP endpoint
                         └── logs ────▶ Grafana Cloud OTLP endpoint

Examining the signals in Grafana

Different questions are answered by each signal:

SignalQuestion
MetricsWhen was an increase in request duration recorded?
TracesWhere was time spent during a slow request?
LogsWhich error was recorded during that request?

A slow period can be identified from a metric graph. The slow operation can be found in a trace, and its error message can be found in a log with the same trace ID. Links between traces and logs are created in Grafana when correlation is configured between the data sources and shared identifiers are included in the records.

Managing Collector configuration across machines

A local YAML file can be used for one Collector. With many machines, a configuration is needed for each Collector. To change an exporter or processor by hand, each file must be found and edited, and each Collector must be restarted. As more machines are added, configuration becomes harder to manage.

Agent configuration and status are carried by OpAMP, the Open Agent Management Protocol. The OpAMP Supervisor is run beside the Collector. A connection to an OpAMP server is established by the Supervisor, and the Collector is started or restarted with the configuration received. The Supervisor is kept outside the telemetry data path.

Separate paths are used for telemetry and management:

Telemetry data:
Application ──OTLP──► Collector ──OTLP/HTTP──► Grafana Cloud

Management:
Grafana Fleet ──OpAMP──► Supervisor ──configures──► Collector

Telemetry continues to be sent by the application to the same local Collector address while configuration is managed through the Supervisor.

Bootstrapping the Supervisor

A local bootstrap is needed before remote configuration can be received. The Collector binary, Supervisor binary, and a small Supervisor configuration must be installed on the host. The OpAMP server address is set in that configuration. The Supervisor can be started by a process manager when the host starts.

A shortened supervisor.yaml is shown below:

server:
  endpoint: wss://<fleet-endpoint>/v1/opamp
  headers:
    Authorization: "Basic <token>"

capabilities:
  accepts_remote_config: true
  reports_effective_config: true

agent:
  executable: /path/to/otelcol-contrib

storage:
  directory: /path/to/supervisor-storage

Placeholders are used for the endpoint, token, and file paths. In a deployment, the endpoint is supplied by Fleet, and credentials are supplied during bootstrap. The agent.executable path is set to the installed Collector binary. Remote configuration is accepted when accepts_remote_config is enabled. An instance UID is used by OpAMP to identify the Collector instance. Matching attributes are used by Fleet to assign configuration pipelines.

The management server address is set in the local file. The Collector pipeline can then be supplied remotely by Fleet.

A description of the Collector is needed by Fleet before a remote configuration can be chosen. A configuration is also needed before the Collector can be started. To resolve this, the Collector can be started by the Supervisor with a temporary no-op configuration. The Collector description is then reported through its OpAMP extension. When remote configuration is received, the Collector is restarted with it. The last known configuration is cached for later restarts if the management server is unavailable.

Supervisor starts
    ↓
Temporary Collector configuration reports identity
    ↓
Fleet sends assigned configuration
    ↓
Supervisor starts Collector with the three pipelines

Sending Collector configuration from Fleet Management

A group of Collectors can be managed through Grafana Fleet Management. After a Supervisor is connected, its Collector is listed in Fleet Inventory. A configuration pipeline can be assigned to matching Collectors. For example, one pipeline can be assigned to Collectors marked environment=staging, while another can be assigned to production Collectors.

The three-signal Collector configuration shown earlier can be assigned through Fleet. The remote configuration is sent over OpAMP, combined with any local configuration by the Supervisor, and used to start the Collector. When the assigned configuration is changed, the Collector is restarted with the updated configuration. The same receiver address continues to be used by the application.

Checking what the Collector actually runs

The configuration offered by Fleet is called the remote configuration. After local and remote inputs are combined by the Supervisor, the configuration used by the Collector is called the effective configuration. Differences can result from that merge.

The effective configuration can be inspected when a change appears in Fleet but different Collector behavior is observed. It shows the configuration given to the Collector by the Supervisor. It can also be reported back to the OpAMP server.

The stages can be checked in order:

  1. A span is found in the Collector’s debug output.
  2. A trace is opened in Grafana.
  3. The request counter is queried as a metric.
  4. A log is found for the same service and trace.
  5. The effective Collector configuration is compared with the configuration assigned in Fleet.