How to Handle Distributed Tracing Across Go Microservices

软件 工程师 都柏林

软件 工程师 都柏林

工程师 精心 打造 可扩展

超越 代码 本身

他们 怎么说

一起 构建

消息 已收到

隐私 政策

使用 条款

Cookie 政策

免责声明

最新 动态

精选 作品

作品 展示

我们 提供 什么

服务 的行业

屏幕 之外

共同 构建

View View
Nben Malla
Nben Malla

Software Engineer

Nben Malla is a software engineer based in Dublin, Ireland, specializing in microservices architecture, legacy system modernization and full stack development.

With experience across FinTech and SaaS, he has delivered scalable backend systems for global clients using Go, Java, Python, Django and Laravel, collaborating with teams across Nepal, Ireland, the Netherlands, New Zealand and the United States.

From leading legacy modernization for global banking clients to architecting microservices and distributed systems, the focus has always been to understand the problem deeply, build it right and deliver software that lasts.

  • 阅读文章 阅读文章

    Tutorials 8 mins

    How to Handle Distributed Tracing Across Go Microservices

    Nben Malla 09 Oct, 2026 8 mins

    How to Handle Distributed Tracing Across Go Microservices

    A single payment request on the Standard Chartered migration touched seven Go services, one message queue, and two external APIs. When it took four seconds instead of four hundred milliseconds, nobody could say which hop caused the delay. Every service produced logs, and none of them connected to each other.

    Distributed tracing solves that problem, but most teams I have seen adopt it badly. They install an SDK, wrap one service, and ship a dashboard that shows traces covering a single hop. Then the traces break at the first queue, and engineers stop trusting the tool.

    This is the approach I use now with OpenTelemetry in Go. It starts with context propagation, not dashboards, and it adds each piece only after the previous one works.

    Start With Context Propagation, Not Dashboards

    A trace is a tree of spans that share one trace ID across services. The trace ID survives only if every service passes it along, and that mechanism is called context propagation. When propagation fails at one hop, the trace splits in two, and you debug two halves instead of one request.

    I configure propagation before I configure anything else. I use the W3C Trace Context format, because every major vendor and every major language supports it. Mixing formats across services produces orphaned spans faster than any other mistake.

    go
    func InitTracer(ctx context.Context, service string) (func(context.Context) error, error) {
        exp, err := otlptracegrpc.New(ctx)
        if err != nil {
            return nil, fmt.Errorf("create otlp exporter: %w", err)
        }
    
        res, err := resource.New(ctx,
            resource.WithAttributes(semconv.ServiceName(service)),
        )
        if err != nil {
            return nil, fmt.Errorf("create resource: %w", err)
        }
    
        tp := sdktrace.NewTracerProvider(
            sdktrace.WithBatcher(exp),
            sdktrace.WithResource(res),
        )
    
        otel.SetTracerProvider(tp)
        otel.SetTextMapPropagator(propagation.NewCompositeTextMapPropagator(
            propagation.TraceContext{},
            propagation.Baggage{},
        ))
    
        return tp.Shutdown, nil
    }

    Every service calls this function once at startup and defers the returned shutdown function. I keep it in a shared internal package so no team writes its own version. Consistency here matters more than any configuration option.

    Instrument the Edges First

    Libraries instrument the boundaries, so I start there. The otelhttp and otelgrpc packages create spans and propagate headers without custom code inside handlers. Wrapping the server and the client in each service gives me full request paths across every synchronous call in an afternoon.

    go
    // Server side: wrap the router once
    handler := otelhttp.NewHandler(mux, "orders-api")
    http.ListenAndServe(":8080", handler)
    
    // Client side: wrap the transport once
    client := &http.Client{
        Transport: otelhttp.NewTransport(http.DefaultTransport),
    }
    
    // gRPC server and client
    srv := grpc.NewServer(grpc.StatsHandler(otelgrpc.NewServerHandler()))
    conn, err := grpc.NewClient(addr,
        grpc.WithStatsHandler(otelgrpc.NewClientHandler()),
        grpc.WithTransportCredentials(creds),
    )

    One rule breaks all of this: always pass the request context. A call built with http.NewRequest instead of http.NewRequestWithContext drops the trace ID, and the downstream span starts a new trace. I found this bug in three services during a single review, and each one had silently split its traces for months.

    I added a lint rule that flags outbound requests built without a context. It catches the mistake in CI instead of during a production investigation. A rule that runs on every pull request costs far less than a debugging session where half the trace is missing.

    Carry Context Through Message Queues

    Queues are where traces die. HTTP and gRPC libraries inject headers automatically, but a message broker client knows nothing about OpenTelemetry. If a producer publishes without injecting, the consumer starts a fresh trace, and the link between cause and effect disappears.

    I solve this with a small carrier type that adapts message headers to the propagator interface. The producer injects the current context, and the consumer extracts it before starting a span. I set the producer and consumer span kinds so backends draw the relationship correctly.

    go
    type headerCarrier []kafka.Header
    
    func (c *headerCarrier) Get(key string) string {
        for _, h := range *c {
            if h.Key == key {
                return string(h.Value)
            }
        }
        return ""
    }
    
    func (c *headerCarrier) Set(key, val string) {
        for i, h := range *c {
            if h.Key == key {
                (*c)[i].Value = []byte(val)
                return
            }
        }
        *c = append(*c, kafka.Header{Key: key, Value: []byte(val)})
    }
    
    func (c *headerCarrier) Keys() []string {
        keys := make([]string, 0, len(*c))
        for _, h := range *c {
            keys = append(keys, h.Key)
        }
        return keys
    }

    Become a Sponsor

    Partner with us as a sponsor and help support our mission while connecting your brand with our community. We offer valuable opportunities to showcase your organization and build meaningful partnerships.

    go
    // Producer: inject before publishing
    carrier := headerCarrier(msg.Headers)
    otel.GetTextMapPropagator().Inject(ctx, &carrier)
    msg.Headers = carrier
    
    // Consumer: extract before starting the span
    carrier := headerCarrier(msg.Headers)
    ctx := otel.GetTextMapPropagator().Extract(context.Background(), &carrier)
    ctx, span := tracer.Start(ctx, "process order event",
        trace.WithSpanKind(trace.SpanKindConsumer))
    defer span.End()

    I wrap this logic in one publish helper and one consume helper per broker, so no handler touches propagation directly. One wrapper, tested once, replaces dozens of fragile copies. This single change made traces useful across the systems I maintain.

    Add Spans Where Decisions Happen

    Automatic instrumentation gives me the skeleton, and I add spans by hand only where the code makes a decision or calls something slow. A span around every function produces noise and inflates storage cost. A span around a charge, a cache lookup, or a rule evaluation tells me where the time went.

    go
    func (s *Service) Charge(ctx context.Context, o Order) error {
        ctx, span := tracer.Start(ctx, "charge order")
        defer span.End()
    
        span.SetAttributes(
            attribute.String("order.id", o.ID),
            attribute.String("payment.provider", o.Provider),
        )
    
        if err := s.gateway.Charge(ctx, o); err != nil {
            span.RecordError(err)
            span.SetStatus(codes.Error, "charge failed")
            return fmt.Errorf("charge order %s: %w", o.ID, err)
        }
        return nil
    }

    Attributes turn a span into evidence. I attach identifiers that engineers search by, such as the order ID and the provider. In financial systems I never attach card numbers, account numbers, or personal data, because traces flow to backends with weaker access controls than the database.

    I also record errors on the span and set its status. Without a status, the backend cannot filter failed traces, and the most valuable traces stay buried among thousands of successful ones.

    Sample Before the Bill Arrives

    Tracing every request at volume costs real money in storage and network. Sampling reduces that volume, and the place where I sample matters. Head sampling decides at the first span, which keeps the cost low but cannot know whether the request will fail.

    go
    tp := sdktrace.NewTracerProvider(
        sdktrace.WithBatcher(exp),
        sdktrace.WithResource(res),
        sdktrace.WithSampler(
            sdktrace.ParentBased(sdktrace.TraceIDRatioBased(0.1)),
        ),
    )

    The parent based sampler makes every service honor the decision made at the edge, so a trace is either complete or absent. Without it, each service samples on its own and produces partial traces that explain nothing. I set a low ratio at the edge, usually ten percent, and adjust it when traffic changes.

    For errors and slow requests, tail sampling at the OpenTelemetry Collector decides after the trace completes. I keep every trace with an error status and a share of the rest. Tail sampling costs collector memory, so I size the collector deliberately and monitor it like any other service.

    yaml
    processors:
      tail_sampling:
        decision_wait: 10s
        policies:
          - name: keep errors
            type: status_code
            status_code:
              status_codes: [ERROR]
          - name: sample the rest
            type: probabilistic
            probabilistic:
              sampling_percentage: 10

    Connect Traces to Logs

    A trace shows where the time went, and logs explain why. The two help each other only when they share an identifier. I add the trace ID and span ID to every log line written inside a request.

    go
    func LoggerFor(ctx context.Context, l *slog.Logger) *slog.Logger {
        sc := trace.SpanContextFromContext(ctx)
        if !sc.IsValid() {
            return l
        }
        return l.With(
            "trace_id", sc.TraceID().String(),
            "span_id", sc.SpanID().String(),
        )
    }

    With this in place, I jump from a slow span to its exact log lines in seconds, and from an error log to its full trace. Before I added it, I searched logs by timestamp and guessed which lines belonged together. That guessing cost hours during incidents.

    I attach the identifiers in one helper, and handlers call it at the start of each request. Engineers adopt a practice when it costs one line. They skip it when it costs a pull request of boilerplate.

    Conclusion

    Distributed tracing does not require a large platform project. It requires propagation that never breaks, instrumentation at the boundaries, and restraint about everything else. Teams that start with dashboards usually discover too late that the data underneath them has holes.

    Each piece in this article exists because a specific failure taught me it was necessary: orphaned spans from a missing context, a queue that split every trace, a storage bill that grew faster than the traffic. I add the pieces in order and verify each one before I start the next.

    A trace you trust is worth more than a hundred dashboards you do not. Engineers open the first kind during an incident, and they ignore the second.