Instrument a GraphQL API: Apollo & Yoga N+1 Detection
Resolver spans + N+1 detection for Apollo Server and GraphQL Yoga. Full OpenTelemetry setup.
Last month I watched a trace waterfall where a single users { posts { comments } } query spawned 847 database calls. The GraphQL layer looked fine — total request time was 2.3 seconds, nothing screaming at me. But the span breakdown told a different story: one resolver running 423 times, each firing two Prisma queries.
Classic N+1. The kind of thing that makes you feel stupid once you see it. Nobody knew it was happening until the trace made it obvious.
This post covers how to instrument a GraphQL API — getting resolver-level spans out of Apollo Server and GraphQL Yoga, detecting N+1 patterns automatically, and the gotchas that'll bite you if you just copy-paste from the docs. By the end you'll have working code that surfaces slow fields and warns you before your database connection pool catches fire.
What We're Building
A fully instrumented GraphQL API with:
- Request-level spans — total query execution time, operation name, errors
- Resolver-level spans — individual field execution times nested under the request span
- N+1 detection — automatic warnings when a resolver fires suspiciously often
- DataLoader integration — batch visibility so you can verify your fixes work
The traces export to JustAnalytics via OpenTelemetry, but the instrumentation itself is standard OTel — swap the exporter if you're using something else. If you're consolidating your observability stack, our guide to replacing GA4, Sentry, and Pingdom covers the full migration.
Prerequisites
- Node.js 18+ (20+ recommended for native fetch)
- Apollo Server 4.x or GraphQL Yoga 5.x
- Basic familiarity with OpenTelemetry concepts (spans, traces, context)
- An existing GraphQL schema — I'll use a simple User/Post/Comment example
- JustAnalytics account with an API key, or any OTel-compatible backend (see our Next.js 15 analytics tutorial for a quick-start example)
Step 1: Install the OpenTelemetry SDK
Start with the core packages:
npm install @opentelemetry/sdk-node @opentelemetry/api \
@opentelemetry/exporter-trace-otlp-http \
@opentelemetry/instrumentation-http \
@opentelemetry/semantic-conventions
The @opentelemetry/sdk-node package handles the boring setup — resource detection, span processors, exporter wiring. Don't use the individual packages unless you need fine-grained control. I wasted two days debugging context propagation issues because I was hand-rolling the SDK setup. Two days. On context propagation. The kind of yak-shaving that makes you question your career choices. Just use the Node SDK.
Step 2: Initialize Tracing Before Your Server Starts
Create a tracing.ts file that runs before anything else imports your GraphQL server:
// tracing.ts
import { NodeSDK } from "@opentelemetry/sdk-node";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
import { HttpInstrumentation } from "@opentelemetry/instrumentation-http";
import { Resource } from "@opentelemetry/resources";
import { ATTR_SERVICE_NAME, ATTR_SERVICE_VERSION } from "@opentelemetry/semantic-conventions";
const sdk = new NodeSDK({
resource: new Resource({
[ATTR_SERVICE_NAME]: "graphql-api",
[ATTR_SERVICE_VERSION]: process.env.npm_package_version || "0.0.0",
}),
traceExporter: new OTLPTraceExporter({
url: "https://otlp.justanalytics.app/v1/traces",
headers: {
Authorization: `Bearer ${process.env.JA_API_KEY}`,
},
}),
instrumentations: [new HttpInstrumentation()],
});
sdk.start();
process.on("SIGTERM", () => {
sdk.shutdown().then(() => process.exit(0));
});
Import this file at the very top of your entrypoint — before Express, before Apollo, before anything that makes HTTP calls:
// index.ts
import "./tracing"; // MUST be first
import { startServer } from "./server";
startServer();
The order matters. OpenTelemetry patches http and https modules at import time. If your server imports before tracing initializes, the patches don't apply and your HTTP spans silently vanish. I've seen this bug three times in production codebases. Twice it was mine.
Step 3: Add Apollo Server Plugin for Resolver Spans
Apollo's plugin API gives you hooks into the request lifecycle. Here's a plugin that creates spans for each resolver:
// plugins/tracing-plugin.ts
import { ApolloServerPlugin } from "@apollo/server";
import { trace, SpanStatusCode, context } from "@opentelemetry/api";
const tracer = trace.getTracer("graphql-resolvers");
export const tracingPlugin: ApolloServerPlugin = {
async requestDidStart({ request }) {
const operationName = request.operationName || "anonymous";
const rootSpan = tracer.startSpan(`graphql.${operationName}`);
return {
async executionDidStart() {
return {
willResolveField({ info }) {
const fieldPath = `${info.parentType.name}.${info.fieldName}`;
const span = tracer.startSpan(
`resolve.${fieldPath}`,
{ attributes: { "graphql.field": fieldPath } },
context.active()
);
return (error) => {
if (error) {
span.setStatus({ code: SpanStatusCode.ERROR, message: error.message });
span.recordException(error);
}
span.end();
};
},
};
},
async willSendResponse() {
rootSpan.end();
},
async didEncounterErrors({ errors }) {
rootSpan.setStatus({ code: SpanStatusCode.ERROR });
errors.forEach((err) => rootSpan.recordException(err));
},
};
},
};
Register the plugin when you create your Apollo Server:
import { ApolloServer } from "@apollo/server";
import { tracingPlugin } from "./plugins/tracing-plugin";
const server = new ApolloServer({
typeDefs,
resolvers,
plugins: [tracingPlugin],
});
Now every resolver execution creates its own span. When you view a trace in JustAnalytics (or Datadog, or Jaeger, or whatever), you'll see the resolver hierarchy: Query.users containing User.posts containing Post.comments. The slow field jumps out immediately. For a deeper dive into distributed tracing, see our guide to how tracing context survives async.
One caveat: this creates a lot of spans. A complex query with 50 fields generates 50+ spans. Fine for debugging. Absolutely maddening in production when you're paying per span. I'll show you how to filter in Step 6.
Step 4: GraphQL Yoga Instrumentation (Alternative)
If you're using GraphQL Yoga instead of Apollo, the pattern is similar but uses Envelop plugins:
// yoga-tracing.ts
import { Plugin } from "graphql-yoga";
import { trace, SpanStatusCode, context } from "@opentelemetry/api";
const tracer = trace.getTracer("graphql-yoga-resolvers");
export const yogaTracingPlugin: Plugin = {
onExecute({ args }) {
const operationName = args.operationName || "anonymous";
const rootSpan = tracer.startSpan(`graphql.${operationName}`);
return {
onExecuteDone({ result }) {
if ("errors" in result && result.errors) {
rootSpan.setStatus({ code: SpanStatusCode.ERROR });
result.errors.forEach((err) => rootSpan.recordException(err));
}
rootSpan.end();
},
};
},
onResolverCalled({ info }) {
const fieldPath = `${info.parentType.name}.${info.fieldName}`;
const span = tracer.startSpan(
`resolve.${fieldPath}`,
{ attributes: { "graphql.field": fieldPath } },
context.active()
);
return ({ result }) => {
if (result instanceof Error) {
span.setStatus({ code: SpanStatusCode.ERROR, message: result.message });
span.recordException(result);
}
span.end();
};
},
};
Then add it to your Yoga server:
import { createYoga } from "graphql-yoga";
import { yogaTracingPlugin } from "./yoga-tracing";
const yoga = createYoga({
schema,
plugins: [yogaTracingPlugin],
});
Yoga also ships with useOpenTelemetry() from @graphql-yoga/plugin-opentelemetry, but it only does request-level spans. The resolver-level granularity requires the custom plugin above.
Step 5: Automatic N+1 Detection
Here's where things get useful. Look, I'll be honest — N+1 queries are the silent killer of GraphQL performance, and they're embarrassingly easy to introduce. I've shipped them to production more times than I'd like to admit. One nested resolver without DataLoader and suddenly you're making 500 database calls.
This plugin tracks resolver execution counts and warns you:
// plugins/n-plus-one-detector.ts
import { ApolloServerPlugin } from "@apollo/server";
import { trace } from "@opentelemetry/api";
const N_PLUS_ONE_THRESHOLD = 5;
export const nPlusOneDetector: ApolloServerPlugin = {
async requestDidStart() {
const resolverCounts = new Map<string, number>();
return {
async executionDidStart() {
return {
willResolveField({ info }) {
const fieldPath = `${info.parentType.name}.${info.fieldName}`;
const count = (resolverCounts.get(fieldPath) || 0) + 1;
resolverCounts.set(fieldPath, count);
return () => {
// Check after resolver completes
if (count === N_PLUS_ONE_THRESHOLD) {
const span = trace.getActiveSpan();
span?.addEvent("n_plus_one_warning", {
"graphql.field": fieldPath,
"resolver.count": count,
"message": `Resolver ${fieldPath} called ${count}+ times — possible N+1`,
});
console.warn(
`[N+1 Warning] ${fieldPath} called ${count}+ times in single request`
);
}
};
},
};
},
};
},
};
Add both plugins to your server:
const server = new ApolloServer({
typeDefs,
resolvers,
plugins: [tracingPlugin, nPlusOneDetector],
});
Now when you run a query that triggers N+1 behavior, you'll see a warning in your logs and an event attached to your trace. The trace annotation is the useful part — you can set up alerts in JustAnalytics when n_plus_one_warning events spike after a deploy. You can also correlate errors with funnel drop-offs to see if N+1 performance issues are impacting user conversion. For deeper resolver-level analysis, our tRPC and GraphQL APM deep-dive covers advanced patterns.
Why threshold of 5? Some legitimate patterns fire the same resolver a few times. Five is high enough to ignore normal usage but low enough to catch actual N+1 patterns. Honestly? The "right" threshold is whatever stops the alert fatigue in your codebase. Start at 5. Bump it up when — not if — you're drowning in warnings.
Step 6: Filtering Noisy Spans in Production
Creating a span for every resolver is overkill in production. A users { id name email } query doesn't need three child spans for trivial field access. Save the granularity for resolvers that actually do work.
Here's a filtered version that only traces resolvers with async execution or explicit annotation:
willResolveField({ info }) {
const fieldPath = `${info.parentType.name}.${info.fieldName}`;
// Skip scalar fields that just return a property
const returnType = info.returnType.toString();
const isScalar = ["String", "Int", "Float", "Boolean", "ID"].some(
(t) => returnType === t || returnType === `${t}!`
);
// Skip unless it's a complex type or marked for tracing
const shouldTrace = !isScalar || info.fieldNodes[0]?.directives?.some(
(d) => d.name.value === "trace"
);
if (!shouldTrace) {
return () => {}; // no-op
}
const span = tracer.startSpan(`resolve.${fieldPath}`);
return (error) => {
if (error) span.recordException(error);
span.end();
};
}
This skips User.name, Post.title, and other scalar fields that resolve instantly. You still get spans for User.posts, Post.comments, and anything returning an object or list type — the resolvers where N+1 problems actually live.
For even more control, add a custom @trace directive to your schema and only instrument fields that have it. Useful when you want to zero in on specific slow paths without the noise. (Though fair warning: I've seen teams over-engineer this and end up tracing nothing useful. Keep it simple.)
Common Errors and How to Fix Them
Traces appear but resolvers are missing spans. You're probably importing your server before tracing.ts. The OpenTelemetry context propagation relies on monkey-patching, and if the patches apply after your modules load, child spans won't link to the parent. Fix: make import "./tracing" the first line of your entrypoint.
Cannot read property 'startSpan' of undefined. The tracer isn't initialized. Either the SDK failed to start (check for errors during startup) or you're calling trace.getTracer() before sdk.start() completes. The NodeSDK start is synchronous in recent versions, so this usually means an import order issue.
Resolver spans show 0ms duration. Some GraphQL frameworks resolve synchronously when the resolver just returns a cached value or property access. That's correct — those resolvers genuinely take sub-millisecond time. If you're seeing 0ms on resolvers that should be slow, check that you're awaiting async operations inside the resolver.
N+1 warnings fire on every request. Your threshold might be too low, or you have a schema pattern where the same resolver legitimately fires many times (like resolving edges in a connection type). Either raise the threshold or exclude specific fields in your detection logic.
"Exporter failed to send" errors. Network issues or wrong endpoint URL. For JustAnalytics, the endpoint is https://otlp.justanalytics.app/v1/traces. Make sure your JA_API_KEY environment variable is set. The OTLP exporter fails silently by default. Silently! In 2026! I cannot overstate how much time I've lost to this. Add error logging to the exporter. Future you will thank present you. If you want help debugging slow API endpoints, our guide to finding P99 latency sources covers the investigation workflow.
Next Steps
Now that you've got resolver spans and N+1 detection, you'll probably want to:
Add database query spans. Prisma has first-party OpenTelemetry support via previewFeatures = ["tracing"] in your schema. Enable it and your Prisma queries appear as child spans under the resolver that triggered them. We covered this in our p99 latency debugging guide — same concept applies to GraphQL. (Fair warning: the first time you see the full trace with both resolver AND database spans, you might not like what you find. I certainly didn't.)
Set up alerting on resolver latency. JustAnalytics can alert when specific resolver spans exceed thresholds. Useful for catching regressions before they hit users. You'd configure this in the APM alerting section, filtering to spans where name LIKE 'resolve.%'.
Integrate DataLoader. The N+1 detector tells you where the problem is. DataLoader solves it. It batches multiple resolver calls into a single database query. The DataLoader docs are solid — and once you add it, your N+1 warnings should disappear. Should. Not always.
For teams running GraphQL in production alongside REST APIs, the same tracing approach works for both. Our uptime monitoring for GraphQL APIs guide covers the health-check side — combining endpoint monitoring with resolver-level APM gives you the full picture.
Frequently Asked Questions
How do you add resolver-level spans to Apollo Server?
Use Apollo's plugin API with willResolveField and didResolveField hooks. Start a span in willResolveField, capture the field name and parent type, then end the span in didResolveField. The OpenTelemetry SDK handles trace context propagation automatically once you set up the instrumentation.
How do you detect N+1 queries in GraphQL automatically?
Track resolver execution counts per field during each request using a Map keyed by field path. When a resolver fires more than a configurable threshold — typically 5-10 times per request — log a warning or create a trace annotation. DataLoader solves the N+1 problem, but detection helps you find un-batched resolvers.
Does GraphQL Yoga support the same instrumentation as Apollo?
Yes. Yoga uses the Envelop plugin system with onExecute and onResolverCalled hooks. The pattern is nearly identical: start spans on resolver entry, capture field metadata, end spans on completion. Yoga also ships with a built-in useOpenTelemetry plugin that handles basic request-level tracing out of the box.
What's the performance overhead of resolver-level tracing?
Typically 2-5% latency increase at high span volume. The overhead comes from span creation and context propagation, not from the actual tracing logic. For production, consider sampling or only creating spans for resolvers that exceed a latency threshold. Request-level spans have negligible overhead.
Try JustAnalytics
All-in-one observability in one under-5KB script: cookieless analytics + error tracking + APM + session replay + uptime + structured logs. Replaces GA4 + Sentry + Datadog + Pingdom + LogRocket. Free tier (100K events/mo), Pro $49/month ($39 annual).
Author at JustAnalytics.