API performance testing measures how fast and how reliably an API responds as traffic grows. Load testing checks the traffic you expect, stress testing pushes past it to find the breaking point, spike testing checks sudden surges, and soak testing runs for hours to expose memory leaks and slow degradation.

API Performance and Load Testing

A functional API test asks "does it return the right answer?". A performance test asks "does it still return the right answer, fast enough, when a thousand people ask at once — and for hours?". This guide explains the types of performance tests, the metrics that matter, how to plan a test, the tools testers use, and includes a small load test you can run yourself.

Advertisement

Why APIs Need Performance Testing

Every screen, mobile app and partner integration depends on APIs, so a slow API slows everything built on it. Performance testing finds:

  • How response time changes as users increase.
  • The maximum load the system handles before errors appear.
  • Bottlenecks — slow queries, missing indexes, thread-pool or connection-pool limits.
  • Memory leaks and degradation over long periods.
  • Whether the system recovers after a spike or failure.

Types of Performance Tests

Type Question it answers Load pattern
Load test Does it meet targets at expected and peak traffic? Ramp up to the expected peak, hold
Stress test Where does it break, and how does it fail and recover? Keep increasing beyond peak
Spike test Can it handle a sudden surge (a sale, a news event)? Jump suddenly, then drop
Soak / endurance test Does it degrade over hours (memory leaks, connection exhaustion)? Normal load for a long time
Scalability test Does adding servers add capacity? Repeat load tests at different sizes

Load vs stress in one line: load testing checks behaviour at the traffic you expect; stress testing finds the limit and checks that failure is graceful — clear 503s or 429s, not corrupted data.

The Metrics That Matter

Metric Meaning Why
Response time percentiles (p50, p95, p99) p95 = 95% of requests were faster than this Averages hide slow outliers; SLAs are usually written as p95 or p99
Throughput Requests per second the API completes Capacity
Error rate Share of failed requests (5xx, timeouts) A fast API that fails isn't fast
Concurrency Simultaneous users or connections The load level being tested
Server resources CPU, memory, DB connections, GC pauses Explains why it's slow

Why percentiles, not averages: if 95 requests take 100 ms and 5 take 5 seconds, the average is about 345 ms — which looks acceptable — while 1 in 20 users waits 5 seconds. The p99 shows that problem immediately.

Planning a Performance Test

  1. Set targets. For example: "Login API — p95 under 300 ms and error rate under 0.1% at 500 concurrent users."
  2. Model the workload. Which endpoints, in what mix (70% search, 20% view, 10% checkout), with realistic think time and data.
  3. Prepare the environment. As close to production as possible, isolated, with monitoring enabled. Never load-test production or third-party APIs without permission.
  4. Prepare data. Enough unique users, products and tokens — reusing one record makes caching hide real behaviour.
  5. Run in steps. Baseline with a few users, then ramp up; hold each level long enough to be stable.
  6. Analyse. Compare percentiles and errors per level against targets, and correlate slow periods with server metrics.
  7. Report and retest after fixes, using the same script and environment.

A Runnable Load Test (Node.js)

This script uses the open-source autocannon package to run three load levels against an endpoint and compare the tail latency with a target. Install with npm install autocannon, set TARGET, and run node load-test.js.

// load-test.js — step through three load levels and check latency against a target
const autocannon = require('autocannon');

const URL = process.env.TARGET || 'http://127.0.0.1:8767/ping';
const levels = [10, 50, 100];            // concurrent connections
const SLA_P95_MS = 200;

(async () => {
  for (const connections of levels) {
    const r = await autocannon({ url: URL, connections, duration: 5 });
    const errors = r.errors + r.non2xx;
    const p97 = r.latency.p97_5;         // autocannon reports p97.5 (no p95); use it as a slightly stricter check
    console.log(
      `${String(connections).padStart(3)} users | ${Math.round(r.requests.average)} req/s | ` +
      `p50 ${r.latency.p50} ms | p97.5 ${p97} ms | p99 ${r.latency.p99} ms | errors ${errors} | ` +
      (p97 <= SLA_P95_MS && errors === 0 ? 'PASS' : 'FAIL'));
  }
})();

Output from a run against a small local test API:

10 users | 228 req/s | p50 43 ms | p97.5 45 ms | p99 45 ms | errors 0 | PASS
50 users | 1042 req/s | p50 44 ms | p97.5 47 ms | p99 52 ms | errors 0 | PASS
100 users | 1085 req/s | p50 43 ms | p97.5 45 ms | p99 47 ms | errors 0 | PASS

Reading it: throughput rose from 10 to 50 users and then flattened at 100 while latency stayed stable — this small service reached its capacity around 1,000 requests per second. With a real API you'd keep increasing until latency climbs or errors appear; that point is the capacity limit. Absolute numbers depend entirely on the machine and network.

Tools

Tool Good for
Apache JMeter The most common enterprise tool: GUI to build test plans, CLI to run them, many protocols
k6 Load tests written in JavaScript, thresholds that fail CI builds, developer-friendly
Gatling Code-based scenarios (Java/Scala/Kotlin) with detailed HTML reports
Locust Python-based user scenarios
Postman Response-time assertions in functional tests; its performance runner suits small, quick checks, not full-scale load
Newman Running Postman collections in CI — a functional tool, not a load generator

Run JMeter tests from the command line for real load (the GUI is for building tests):

jmeter -n -t login-load.jmx -l results.jtl -e -o report/

Response-time checks in functional tests

A cheap early warning in your normal suites — not a replacement for load testing:

// Postman Tests tab
pm.test('Login responds in under 800 ms', () => pm.expect(pm.response.responseTime).to.be.below(800));
// REST Assured
given().body(credentials).when().post("/login").then().statusCode(200).time(lessThan(800L));

More: Postman Scripting and REST Assured Response Validation.

Common Bottlenecks

  • Slow database queries or missing indexes.
  • Connection-pool or thread-pool limits reached.
  • N+1 calls — one request triggering many downstream calls.
  • No caching for data that rarely changes.
  • Large payloads without pagination or compression.
  • Synchronous calls to slow third-party services.

Design the functional and negative checks for the same endpoints in the API Testing Scenario Lab. Broader performance concepts: How to Do Performance Testing.

Interview Answer

"For API performance testing I agree targets first — for example p95 under 300 ms and error rate under 0.1% at expected peak — then build a realistic workload in JMeter, run a baseline and ramp up in steps, and track percentiles, throughput and error rate alongside server metrics. Load testing checks expected and peak traffic; stress testing pushes past it to find the breaking point and check graceful failure. In our functional suites we also assert response times as an early warning."

From Real Projects

My API testing experience covers CRUD operations, HTTP methods, JSON path, and validating JSON and XML data. On Testsigma, a platform that itself automates API tests, understanding requests, responses and status codes was part of understanding the product. Whatever tool you use, validate the status code, the response structure and the values — not just that the call returned something. Measure response-time percentiles, not just averages.

📚 Official documentation: Apache JMeter user manual · Grafana k6 documentation

API Stress Testing: How to Find the Breaking Point

A load test answers "are we fine at normal traffic?". A stress test answers "how much can we take, and what fails first?". A simple, repeatable approach:

  1. Start from a baseline at expected traffic and note p95 response time and error rate.
  2. Increase load in steps, for example 25% more virtual users every few minutes, rather than all at once.
  3. Watch the API and its dependencies: p95/p99 latency, error rate, CPU, memory, database connections and queue depth.
  4. Define "broken" before you start, for example error rate above 1% or p95 above your SLA.
  5. Record the highest load the API handled within those limits; that's your capacity.
  6. Drop the load and check recovery. A healthy API returns to baseline on its own; one that stays slow has a leak or a stuck resource.

The same idea as a k6 script, where stages ramp the users up and thresholds fail the run automatically:

import http from 'k6/http';
import { sleep } from 'k6';

export const options = {
  stages: [
    { duration: '2m', target: 50 },    // normal load
    { duration: '5m', target: 200 },   // above normal
    { duration: '5m', target: 400 },   // stress
    { duration: '2m', target: 0 },     // recovery
  ],
  thresholds: {
    http_req_failed: ['rate<0.01'],    // under 1% errors
    http_req_duration: ['p(95)<800'],  // 95% of requests under 800 ms
  },
};

export default function () {
  http.get('https://test-api.example.com/products');
  sleep(1);
}

Run stress tests against a production-like test environment, never against production without agreement, because the goal is to break things.

FAQs

Why is performance testing important for APIs?

Every client depends on the API, so its speed, capacity and stability under load decide the experience for all of them.

What is the difference between load and stress testing?

Load testing checks performance at expected and peak traffic; stress testing goes beyond peak to find the breaking point and check how the system fails and recovers.

What is p95 response time?

The time within which 95% of requests completed. It's preferred over the average because it shows what slower users experience.

Can Postman be used for load testing?

Only for small checks. Use JMeter, k6, Gatling or Locust for realistic load; Postman and Newman are best for functional tests with response-time assertions.

Which tools are used for API performance testing?

JMeter most commonly, plus k6, Gatling and Locust — with monitoring tools to see server-side resource use.

Should performance tests run against production?

Generally no — use a production-like, isolated environment. Production testing needs explicit approval, careful limits and monitoring.

What is API stress testing?

API stress testing sends more traffic than the API is designed for, increasing it step by step, to find the point where response times or errors become unacceptable and to check that the API recovers once the load drops.

How many virtual users should an API load test use?

Base it on real traffic: take your peak requests per second from production logs or analytics, then size the test to match it for a load test, and to two or three times it for a stress test. Guessing a round number like 1,000 users tells you very little.

What is a good API response time?

It depends on the API, but a common target for user-facing REST endpoints is a p95 under 300 to 800 milliseconds. Measure percentiles, not averages: an average can look fine while 5% of users wait several seconds.