FlowingDev

HAR files, explained: your browser's black box recorder

Learn what HAR (HTTP Archive) files are, how they capture every network request your browser makes, and why they're essential for debugging web performance.

Try the tool: HAR Viewer

In one sentence

A HAR file is a JSON-formatted log of a web browser's interaction with a site, capturing every single network request and response in minute detail for later analysis.

The problem it solves

Picture this: a user on the other side of the world DMs you, "Your app is unusably slow." You try it. It's snappy. They say it's broken. You say it works on your machine. Stalemate.

This is the classic he-said-she-said of web development. Before modern browser tools, debugging remote network issues was a nightmare of guesswork, server log spelunking, and asking non-technical users to describe esoteric error messages. Even with the advent of browser DevTools and their glorious Network tab, the problem remained: the data was ephemeral. You couldn't easily bottle it up and send it to a colleague. A screenshot of a waterfall chart doesn't tell the whole story.

Enter the HTTP Archive format, or HAR. Conceived by the Web Performance Working Group of the W3C, it was designed to be a standard, shareable format for, well, archiving HTTP transactions. It's the digital equivalent of putting a black box recorder in a user's browser.

A HAR file solves the "it works on my machine" problem by capturing the entire network conversation between a browser and a server for a given web page load. It records every request for an image, a script, a font, or an API call. It logs the exact headers sent, the cookies exchanged, the redirects followed, and, most importantly, the precise timings for every stage of the request.

This lets a developer in San Francisco see exactly what a user in Singapore experienced, millisecond by millisecond, without having to guess. It makes transient, hard-to-reproduce network bugs analyzable and turns vague complaints of "slowness" into actionable data.

How it works under the hood

At its core, a HAR file isn't magic. It's just a big, structured JSON file. You can open one in a text editor and see everything, though a dedicated viewer makes it infinitely easier to parse. Let's peek inside.

The Grand Structure: It's Just JSON

A HAR file contains a single top-level JSON object with one key: log. Everything else lives inside this log object.

{
  "log": {
    "version": "1.2",
    "creator": { "name": "Chrome", "version": "118.0.0.0" },
    "browser": { "name": "Chrome", "version": "118.0.0.0" },
    "pages": [ /* ... one or more page objects ... */ ],
    "entries": [ /* ... one or more request/response objects ... */ ]
  }
}
  • version: The HAR spec version, usually "1.2".
  • creator / browser: Metadata about what tool and browser generated the file. Useful for context.
  • pages: An array describing the main page or pages that were loaded. It includes the page title and timings for high-level events like onLoad and onContentLoad.
  • entries: This is the star of the show. It's a long array where each object represents a single network request and its corresponding response.

The Star of the Show: The entries Array

When you analyze a HAR file, you spend 99% of your time in the entries. Each entry is a complete dossier on one resource.

Here's a simplified look at a single entry:

{
  "startedDateTime": "2023-10-27T10:30:05.123Z",
  "time": 258.45,
  "request": { /* ... details of the request ... */ },
  "response": { /* ... details of the response ... */ },
  "timings": { /* ... the juicy performance breakdown ... */ },
  "pageref": "page_1"
}
  • startedDateTime: The exact UTC timestamp when the request started.
  • time: The total elapsed time for the request in milliseconds, from start to finish.
  • request: An object containing everything the browser sent to the server.
  • response: An object containing everything the server sent back.
  • timings: The goldmine for performance debugging. We'll dissect this next.
  • pageref: An ID that links this request back to one of the pages in the pages array.

Anatomy of a Request and Response

The request and response objects are mirrors of what you'd see in DevTools.

The request object details:

  • method: GET, POST, PUT, etc.
  • url: The full URL of the resource.
  • headers: An array of all request headers, like User-Agent, Accept, and Cookie.
  • queryString: An array of any query parameters on the URL.
  • postData: For POST requests, this holds the payload, like form data or a JSON body.

The response object details:

  • status: The HTTP status code (e.g., 200, 404, 500).
  • statusText: The reason phrase (e.g., OK, Not Found).
  • headers: An array of all response headers, like Content-Type, Cache-Control, and Set-Cookie.
  • content: An object describing the response body, including its size, mimeType, and often the body itself in the text property (though this can be omitted to save space or for security).

The Timings Waterfall Breakdown

The timings object is what powers the colorful waterfall chart in a HAR viewer. It breaks down the total request time into its constituent phases. Understanding these is key to diagnosing "why" a request was slow.

Timing What it means
blocked Time the request spent waiting in the browser's queue before it could even start. Often due to connection limits.
dns Time spent on DNS lookup. A high value might indicate a slow DNS provider.
connect Time taken to establish a TCP connection to the server. Includes ssl time.
ssl (Part of connect) Time for the SSL/TLS handshake. High values can point to server config or network issues.
send Time spent sending the HTTP request to the server. Usually very short.
wait Time To First Byte (TTFB). The critical one. Time spent waiting for the server to process the request and send the first byte of the response. A long wait time is almost always a backend problem.
receive Time spent downloading the response body from the server. A long receive time on a small file could indicate a slow network; on a large file, it's expected.

The total time for an entry is the sum of these individual (non-negative) timings. When a HAR viewer shows you a bar for a request, it's visually stacking these timings values end-to-end.

Real-world stories

The Case of the Mysterious Slowness

A frantic PM messages the team: "The new checkout page is super slow for our biggest client! They're threatening to leave!" The dev team tries the checkout flow. It's lightning-fast. The client insists it takes 20 seconds to confirm an order. Instead of a fruitless back-and-forth, the lead dev walks the client through exporting a HAR file.

Upon opening the file, the problem is instantly obvious. In the entries, the POST request to /api/v1/finalize_order has a total time of 20,145ms. Looking at the timings object, the wait (TTFB) is over 20,000ms. The backend server is taking 20 seconds to respond. It turns out this specific client had a massive order history, and an unoptimized database query was timing out, but only for their account. The HAR file provided the smoking gun that pointed directly to a specific backend process.

The Lesson: A HAR file captures user-specific conditions (like account data) that you can't replicate, turning a mystery into a targeted bug report.

The Bloated Bundle Culprit

A marketing site goes live and the bounce rate is through the roof. It just feels heavy. A front-end dev pulls up the site, opens DevTools, records a session, and exports the HAR.

In the HAR viewer, they sort the entries by size. At the top is main.acb123.js at a whopping 5.2 MB. The waterfall shows it's a render-blocking resource; nothing appears on the page until this behemoth has finished downloading. The receive time alone is multiple seconds, even on a fast connection. Worse, looking at the response headers for this entry, they see the server isn't sending a Content-Encoding: gzip header, despite the browser sending Accept-Encoding: gzip in the request headers. The JavaScript bundle wasn't being compressed.

The Lesson: HAR files make it trivially easy to spot performance killers like oversized assets and server misconfigurations that are murdering your page load time.

The Infinite Redirect Loop

A user complains they can't log in. They enter their credentials, click "Log In," and are immediately kicked back to the login page with no error. It's a classic loop. Support asks the user for a HAR file of the login attempt.

The entries list in the HAR tells a clear story:

  1. POST /login succeeds and gets a 302 Redirect to /dashboard. The response includes a Set-Cookie header with the session token.
  2. The browser follows the redirect and makes a GET /dashboard request.
  3. The server responds to GET /dashboard with a 302 Redirect back to /login.

Why? The dev inspects the GET /dashboard request entry. The Cookie header is missing the session token. Then they check the response from the initial POST /login. The Set-Cookie header was session_id=...; Secure; HttpOnly. The Secure flag means the browser will only send the cookie over HTTPS. The user was on a http://staging.example.com environment. The browser was correctly refusing to send the secure cookie over an insecure connection, so the server never saw them as logged in.

The Lesson: HAR files give you a perfect, frame-by-frame replay of HTTP redirects and header exchanges, making it possible to debug complex authentication flows that fail silently.

Common mistakes and traps

  • Forgetting to "Preserve log". If your bug involves moving from page A to page B, you must enable the "Preserve log" (or equivalent) option in DevTools. Otherwise, the log is cleared on navigation, and your HAR file will only contain requests for page B.
  • Sharing sensitive data. HAR files are indiscriminate recorders. They will capture API keys, session tokens in cookies, and personally identifiable information in POST bodies. Always sanitize HAR files before sharing them in public bug trackers or forums.
  • Misinterpreting blocked time. A high blocked time doesn't always mean the network is congested. Browsers have a limit on how many parallel connections they'll open to a single domain (usually 6). If you fire off 20 image requests at once, 14 of them will sit in a blocked state, waiting for one of the first 6 to finish.
  • Ignoring the cache state. If you're testing first-load performance, you need to record with the browser cache disabled. Otherwise, you'll see a lot of 304 Not Modified responses or requests that complete in under a millisecond, which doesn't reflect a new user's experience.

Why it belongs on your radar

You should think about using a HAR file whenever network communication is a potential suspect.

  • When a user reports a performance issue you can't reproduce.
  • When you need to optimize a slow-loading page and want to identify the biggest bottlenecks.
  • When you're debugging a multi-step API flow, like an OAuth login or a payment process, and need to see the exact sequence of events.
  • When you need to file a bug report with a third-party service (like a CDN or API provider) and want to provide them with irrefutable proof of the issue. A HAR file is the universal language of network problems.

Go deeper

Theory done. Time to get your hands dirty — 100% in your browser.

Try the tool: HAR Viewer