APM metrics: complete guide for performance testing teams

Diego Salinas
Enterprise Content Manager
Gatling
Table of contents
Add to Google preferred sources

Summarize this article with AI

APM metrics: complete guide for performance testing teams

APM metrics are the quantifiable measurements that track your application's health, speed, and efficiency: response times, error rates, throughput, and resource utilization across your entire stack. They're what stand between you and the 3 a.m. phone call about production being down, and by the time that call comes, you're already counting the cost in lost transactions and lost trust.

This guide covers the core metrics every performance testing team should track, how infrastructure, trace, and business metrics fit into the picture, how to actually collect them, and how to connect your load testing results directly to production monitoring. Everything here sits inside the broader observability discipline usually shortened to MELT: metrics, events, logs, and traces. This guide focuses on the metrics piece, but you'll see events, logs, and traces show up throughout, because none of the four work in isolation.

What are APM metrics

APM (Application Performance Monitoring) metrics are quantifiable measurements that track the health, speed, and efficiency of software applications. They focus on four core areas:

  • Response time
  • Error rates
  • Throughput
  • Resource utilization

APM tools collect these measurements continuously across your entire application stack, from frontend interfaces to backend services and underlying infrastructure. The goal is straightforward: spot problems before users do. When response times creep up or error rates spike, APM metrics give you the data to investigate and fix issues quickly.

Why APM metrics matter for performance testing teams

Here's something useful to know: load testing tools and APM platforms track the same core metrics. Response times, throughput, error rates, latency percentiles: they're identical whether you're running a Gatling simulation or monitoring production traffic in Datadog.

That overlap creates a direct connection between testing and production. When your load test shows a p95 latency of 200ms under 1,000 concurrent users, you can compare that number directly against what your APM tool reports in production. If production latency suddenly jumps to 350ms, you have a concrete reference point for investigation.

Without this shared vocabulary, performance testing happens in isolation. Teams run tests, see results, and hope those numbers translate to real-world behavior. With APM metrics as your common language, you can validate assumptions and catch regressions before they reach users.

Essential application performance monitoring metrics to track

Application-layer metrics form the foundation of any monitoring strategy. They measure what your code is actually doing, independent of the servers running it.

Apdex score

Apdex (Application Performance Index) translates raw response times into a standardized satisfaction score between 0 and 1. You define a threshold, say 500ms, and the formula categorizes every response as satisfied, tolerating, or frustrated based on how it compares to that threshold.

The score is particularly useful for communicating with stakeholders who don't want to interpret percentile charts. An Apdex of 0.94 means "most users are happy." An Apdex of 0.67 means "we have a problem." Many teams use Apdex thresholds directly in their SLAs.

SLA scores and how they connect to load testing

An SLA (Service Level Agreement) is the commitment you make externally, often to a customer or a business stakeholder: "99.9% of requests complete under 300ms" or "99.95% uptime per month." Apdex measures satisfaction; an SLA score measures compliance against that specific promise, usually tracked as a percentage of time (or percentage of requests) the commitment held.

The useful move is treating your pre-production tests as a proxy for that SLA before it's ever at risk in production. In Gatling Enterprise, you configure SLOs (Service Level Objectives) as response-time percentile bounds or error-ratio bounds evaluated continuously over a run, then reported as the percentage of time the condition held: green at 99% or above, orange between 90 and 99%, red below 90%. If SLOs are configured on a test, they take over from code-defined assertions for that run entirely, so you get one clear pass/fail signal per objective instead of reconciling two systems.

Map your production SLA directly onto a test SLO: if the business commitment is "99.9% of checkout requests under 300ms," configure that exact bound as an SLO on your checkout load test. A test that consistently comes back green is telling you the SLA is safe under the traffic levels you tested. A test that drifts orange is an early warning, weeks before it would show up as a real SLA breach.

Response time and latency percentiles

Average response time can be misleading. If 95% of your requests complete in 100ms but 5% take 3 seconds, your average might look acceptable while thousands of users experience frustration. This isn't a minor statistical quirk: Gatling's own documentation is explicit that mean and standard deviation only make clean sense on symmetric, single-peaked distributions, and real-world response times are almost never that shape. Two very differently shaped distributions can share the same mean.

Percentiles tell the full story:

  • p50 (median): The typical user experience. Half of all requests are faster than this value.
  • p95: What slower requests look like. Only 5% of users experience worse performance.
  • p99: The worst-case scenarios, excluding extreme outliers. Critical for understanding your most impacted users.

When setting performance goals, p95 and p99 matter more than averages. They reveal the experience of users who might otherwise leave without complaining.

Turn this into something actionable rather than a chart you glance at: set a concrete alert on p95, for example "page on-call if p95 exceeds 400ms for 5 consecutive minutes," and a separate, tighter one on p99 for your most business-critical endpoint (checkout, login, search). Alerting on the median alone will miss exactly the users you most need to catch.

Request rate and throughput

Throughput measures capacity: how many requests your application handles per second (RPS) or per minute (RPM). This metric answers fundamental questions about scale.

Can your checkout service handle 500 transactions per second during a flash sale? What happens when traffic doubles? Throughput trends also reveal problems: a sudden drop might indicate upstream failures, while unexpected spikes could signal bot traffic or a viral moment.

A useful threshold pattern here is relative, not absolute: alert when throughput drops more than 30% below the trailing 7-day average for the same time of day and day of week, rather than a single fixed number that breaks the first time a holiday changes your traffic shape.

Error rate

Error rate tracks failed requests as a percentage of total requests. A 0.1% error rate sounds small until you realize that's 1,000 failures per million requests.

The metric becomes most valuable when correlated with other signals. Low latency with high errors might indicate fast failures: your service is rejecting requests quickly. High latency with rising errors often points to timeouts or resource exhaustion.

A concrete alert worth setting from day one: page on-call if more than 2% of the last 100 requests to a given endpoint fail, using a rolling window rather than a single spike, so one bad request doesn't wake anyone up at 3 a.m. for nothing.

Infrastructure metrics for application performance

Application metrics tell you what's happening. Infrastructure metrics help explain why. When response times spike, these measurements point toward root causes.

CPU and memory utilization

CPU utilization above 80% sustained often indicates a performance bottleneck. Your application might be doing too much work per request, running inefficient algorithms, or simply undersized for current traffic.

Memory pressure creates different symptoms. Gradual increases suggest memory leaks. Sudden spikes might indicate large payload processing or cache misses. When memory runs low, applications start swapping to disk or triggering aggressive garbage collection, both devastating for latency. A common baseline: alert on sustained CPU above 80% for more than 5 minutes, and separately on memory that climbs more than 20% over a 30-minute window without a corresponding traffic increase, since that pattern usually means a leak rather than legitimate load.

Garbage collection metrics

For applications running on managed runtimes like the JVM (Java, Scala, Kotlin), garbage collection directly impacts user experience. During GC pauses, your application literally stops processing requests.

Track GC frequency and duration. Minor collections happening constantly suggest your application creates too many short-lived objects. Major collections taking hundreds of milliseconds will show up as latency spikes in your p99 metrics.

Instance count and node availability

Uptime percentage measures reliability: 99.9% availability still means 8.7 hours of downtime per year. For critical services, even 99.99% might not be enough.

Instance count matters in auto-scaling environments. If your application scales from 3 to 15 instances during peak traffic, that's useful capacity planning data. If it scales to 15 instances and still struggles, you've found a bottleneck that horizontal scaling can't solve.

In Kubernetes or other cloud-native environments, track node availability alongside instance count: how many nodes in the cluster are Ready, how often pods restart, and whether the scheduler is struggling to place new pods (a sign you're hitting resource quotas before you hit application limits). A pod-restart count that climbs during a load test is often the first symptom of a memory leak, well before it shows up in an application-level metric.

Network usage and I/O rates

Two infrastructure signals that are easy to overlook until they're the actual bottleneck:

  • Network usage / bandwidth: Inbound and outbound throughput at the host or container level. A service that looks CPU- and memory-healthy but still degrades under load is frequently saturating a network interface or hitting a cloud provider's bandwidth cap.
  • I/O (disk read/write) rates: Storage throughput, tracked separately from database query time. Log-heavy services, file uploads, or anything writing large payloads to disk can bottleneck here even when the database itself is fine.

Both are worth graphing next to CPU and memory on the same dashboard. A production incident that looks like "the app is just slow" often turns out to be a specific one of these four resources maxed out, and having all four on one screen cuts the investigation from an hour to a few minutes.

APM trace metrics and transaction monitoring

With most organizations now running on microservices, modern applications rarely exist as monoliths. A single user request can touch dozens of interconnected components spanning services, databases, and external APIs. Trace metrics follow that journey.

Distributed trace metrics

A trace captures the complete path of a request through your system. Each step, a service call, a database query, a cache lookup, becomes a span with its own timing data.

When a checkout request takes 2 seconds, traces show you exactly where that time went. Maybe 1.5 seconds happened in a single database query. Maybe latency accumulated across 20 microservice hops. Without traces, you're guessing. With them, you know precisely which component to optimize.

This is also where load testing and APM correlation earns its keep. In a past Gatling Enterprise demo, a test showed a 10% error rate under load with no further detail. Correlating that same time window against Dynatrace traces pointed to a database connection pool exhaustion issue as the actual root cause, not the service the errors were reported against. The load test told the team what failed; the trace told them why.

Database query performance metrics

Slow queries cause more performance problems than almost any other factor. A single unoptimized query running on every request can bring down an entire application.

Key database metrics to watch:

  • Query execution time: Both average and p95, broken down by query type.
  • Connection pool utilization: Running out of connections causes requests to queue.
  • Lock contention: Queries waiting on locks indicate concurrency issues.

Adding an index or rewriting a join often delivers large, immediate latency improvements with minimal code changes.

End user experience monitoring metrics

Server-side metrics capture what your infrastructure experiences. Real User Monitoring (RUM) captures what actual users experience in their browsers, and the two can differ dramatically.

Page load time

A server might respond in 50ms, but the user's browser still takes 3 seconds to render the page. Network latency, asset loading, JavaScript execution, and rendering all add up.

Key components include Time to First Byte (TTFB), First Contentful Paint (FCP), and Largest Contentful Paint (LCP), Google's Core Web Vitals. These metrics often reveal optimization opportunities invisible to backend monitoring: uncompressed images, render-blocking scripts, or CDN misconfigurations.

User session metrics

Session duration, bounce rates, and conversion funnels connect technical performance to business outcomes. A 500ms increase in page load time might correlate with a measurable drop in conversions.

This connection helps prioritize performance work. Optimizing a page that 80% of users visit delivers more value than perfecting a rarely used admin screen.

Business and capacity metrics

A handful of metrics sit at the boundary between engineering and the rest of the business. They're worth tracking alongside the technical ones because they're what actually gets a performance-testing budget renewed.

  • Transaction volume: The count of completed business transactions (orders, logins, bookings) rather than raw requests. Throughput tells you how much traffic your system handled; transaction volume tells you how much business that traffic represented, which matters for capacity planning around known peak events.
  • Cloud spend / cost metrics: What it costs to serve current load, and how that cost scales as traffic grows. Teams that skip this end up discovering, only after a scaling event, that handling 3x traffic meant 5x the infrastructure bill. Several Gatling customers use load testing explicitly for this: identifying over-provisioned services before they become a recurring cost, or validating that autoscaling actually scales back down (not just up) once load subsides.

How to collect and analyze APM metrics

None of the metrics above are useful until they're actually flowing somewhere you can query and alert on them. At a high level, the pipeline looks the same regardless of vendor:

  1. Instrumentation. Either an auto-instrumentation agent bundled with your APM vendor's SDK, or a vendor-neutral approach via OpenTelemetry, which exports both metrics and logs to any OTel-compatible collector (Grafana, Prometheus, and others). OpenTelemetry represents response-time data as exponential histograms, which preserves percentile accuracy without storing every raw value.
  2. Collection and aggregation. Agents or collectors ship data to a backend that aggregates it, computes percentiles, and retains it at whatever resolution your retention policy allows.
  3. Log aggregation. Metrics tell you something is wrong; logs tell you the specifics. Centralizing logs (rather than leaving them on individual hosts) is what makes it possible to correlate a metric spike with the exact error messages happening at that moment.
  4. Alerting. Rules defined against the metrics above, ideally with rolling windows rather than single-point spikes, routed to whoever is on call.
  5. Dashboards. The shared view that turns raw numbers into something a team actually looks at daily, not just during an incident.

The reason this matters for a performance-testing team specifically: if your load tests write into the same pipeline, using the same instrumentation, you get one dashboard that shows both pre-production and production data side by side, instead of two disconnected tools that happen to measure similar things. That's the setup the next section builds on.

How to connect load testing results to APM metrics

Load testing and APM work best together. One validates performance before deployment; the other monitors it afterward. The metrics they share make this partnership possible.

Establishing performance baselines before production

Load tests create controlled conditions for measuring performance. Run a test with 1,000 concurrent users, and you know exactly what your p95 latency looks like at that load level.

These baselines become your reference points. When APM shows p95 latency climbing in production, you can compare against your test results. Is current traffic higher than what you tested? Did a recent deployment change performance characteristics?

Correlating test throughput with production traffic

Effective load tests simulate realistic conditions. If production handles 200 RPS during normal hours and 800 RPS during peaks, your tests can cover both scenarios.

APM data tells you what "realistic" actually means. Pull traffic patterns from your monitoring tools, then replicate those patterns in your load tests.

This approach catches problems that synthetic, steady-state tests miss, like race conditions that only appear during traffic ramps.

Using APM metrics as load test assertions

Gatling separates two mechanisms here, and it's worth knowing which one you're using. Code-defined assertions check a statistic (response time, requests per second, failed or successful request count) against a condition (less than, greater than, between) at a chosen scope: globally, across all requests, or on a specific named request. That's the "fail a build if p95 exceeds 500ms" rule most teams start with.

Gatling Enterprise adds SLOs on top: the same kind of threshold, but evaluated continuously through the run and reported as a percentage of time it held, rather than a single pass/fail at the end. If SLOs are configured, they replace assertions for that run entirely. Either way, ramp-up and ramp-down windows can be excluded from the calculation, so a slow warm-up period doesn't unfairly fail an otherwise healthy test. Stop criteria are the third piece: a run can be terminated early if mean CPU, global error ratio, or response time at a given percentile breaches a threshold mid-test, which saves compute and avoids hammering a system that's already clearly failing.

Gatling integrates directly with APM platforms like Datadog and Dynatrace, tagging test-generated traffic by team, test, scenario, and status so it's distinguishable from real traffic in the same dashboard, and streaming test metrics alongside production data.

What this looks like in practice: LoginRadius, a customer identity and access management platform, moved off a homegrown JMeter framework that topped out around 1,000 RPS and ran manually outside CI/CD. After moving to Gatling Enterprise with tests wired into GitLab pipelines, they reduced p95 API latency from 500ms to under 250ms and cut production performance regressions by more than 80%, by catching them in CI before code shipped rather than after.

As their Director of Product and Engineering put it: "Any team aiming to deliver great service should make performance testing a mandatory part of their SDLC. It shouldn't be an afterthought."

How to choose the right application metrics for your stack

Not every metric matters equally for every application. Your architecture and business requirements determine which measurements deserve attention.

Performance priorities by application type METRICS • GUIDE
Application type Priority metrics
Web applications Page load time, Apdex score, error rate
APIs and microservices Latency percentiles (p95/p99), throughput, distributed trace metrics
Data-intensive apps Database query time, GC metrics, memory utilization
Real-time systems p99 latency, connection metrics, availability

Start with the four golden signals, latency, traffic, errors, and saturation, then add specificity based on what your users care about. An e-commerce site might prioritize checkout latency. A real-time collaboration tool might focus on p99 message delivery times.

Connecting load testing to observability platforms

Load testing becomes significantly more valuable when its metrics flow into your observability stack. Gatling Enterprise Edition supports integrations with major platforms, allowing teams to correlate synthetic load with real infrastructure signals.

Datadog

With the Datadog integration, Gatling pushes 25 or more metrics automatically, request timing, TCP connections, TLS handshakes, bandwidth, tagged by team, test, scenario, and status, plus injection start/end events. You can overlay test windows with infrastructure metrics, helping you identify exactly when latency increased and which components were affected.

Dynatrace

The Dynatrace integration enables correlation between load test traffic and distributed traces. You can tag test-generated requests and analyze them at code level, making microservice bottlenecks visible under synthetic stress, the same mechanism behind the connection-pool example earlier in this guide.

New Relic

With New Relic, you can centralize load testing and APM analysis in one place. Test runs appear alongside production telemetry, making regression comparison straightforward.

InfluxDB

Teams using InfluxDB can push load test metrics into time-series databases and visualize them in Grafana. This is particularly useful for long-term trend analysis and custom dashboards.

OpenTelemetry

OpenTelemetry provides a vendor-neutral way to export metrics and traces. Integrating load testing into OpenTelemetry pipelines ensures your synthetic traffic participates in the same observability architecture as your production systems.

Using APM metrics as CI/CD gates

Performance should not be evaluated manually after deployment, especially as teams move toward CI/CD performance automation.

Modern teams define acceptance criteria directly in their pipelines, turning performance testing into a release gate rather than a reporting exercise. Gatling Enterprise Edition supports run stop criteria and SLO thresholds, and integrates with GitHub Actions, GitLab CI, Jenkins, Azure DevOps, Bamboo, and TeamCity using the same pattern everywhere: trigger a simulation, wait for completion, check assertion or SLO results on the pipeline. Canceling the pipeline stops the Gatling run too, so a cancelled deploy doesn't keep burning test credits in the background.

For example:

  • Fail a build if p95 exceeds 500ms.
  • Stop a test if the global error ratio rises above 2%.
  • Abort execution if injector CPU exceeds safe limits.

A deploy that regresses p95 by 15% fails the quality gate automatically, the same way a broken unit test would.

From monitoring to continuous performance visibility

Catching performance issues in production is reactive. Catching them during load testing is proactive. Catching them inside CI is preventative.

When load testing integrates with your APM system, performance becomes observable across the entire lifecycle. This shift aligns with how larger engineering organizations modernize performance engineering: instead of running isolated load tests, teams build continuous performance visibility, and can track it over time with a 0-100 campaign score that rolls up test methodology, whether objectives were met, and how stable results have been across the last few runs.

Turn APM metrics into continuous performance visibility

APM metrics become most valuable when they're part of a continuous strategy rather than occasional checkups. Catching issues in production is good. Catching them during load testing is better. Catching them in CI/CD before merge is best.

Teams using Gatling can stream load test metrics directly to their APM platforms, creating a single view of performance from development through production. The same dashboards that monitor production can also display test results, making comparisons immediate and obvious.

Explore Gatling Enterprise to see how continuous performance visibility works in practice.

{{card}}

About the author
‍
Diego Salinas
Gatling

Diego Salinas Gardón is a senior technical copywriter and content strategist specializing in performance testing, developer tools, SaaS, and software infrastructure. With hands-on experience in front-end development and modern web technologies.

He currently works at Gatling, where he creates content that helps developers and engineering teams better understand performance, testing, and DevOps.

FAQ

What does APM stand for in application monitoring?

APM stands for Application Performance Monitoring. It's the practice of tracking and optimizing how software applications perform in real time, using metrics, traces, and logs to maintain visibility across the entire application stack.

What is the difference between APM metrics and infrastructure monitoring metrics?

APM metrics focus on application behavior: response times, error rates, and transaction traces. Infrastructure monitoring tracks underlying resources like CPU, memory, network, and disk. Most APM platforms collect both, since application problems often have infrastructure root causes.

What is the difference between APM metrics and load testing metrics?

APM metrics measure production application behavior with real user traffic. Load testing metrics measure behavior under simulated traffic in controlled environments. Both track similar KPIs, response time, throughput, error rate, which makes them complementary tools for comprehensive performance visibility.

Can APM metrics predict performance issues before they affect users?

To a point. APM metrics are inherently reactive: they tell you a problem is starting, not one that hasn't happened yet, though a slow, sustained trend (rising p95, climbing memory, shrinking connection pool headroom) is often visible well before it causes a user-facing incident. The more reliable way to predict issues before they affect users is to reproduce the traffic pattern in a load test first, where a regression shows up as a failed assertion in CI rather than a page at 3 a.m.

Ready to move beyond local tests?

Start building a performance strategy that scales with your business.

Need technical references and tutorials?

Minimal features, for local use only