The honest answer to "how much does eDiscovery cost" is that it depends on how much data you have, how many people touched it, and how many documents a human being ends up reading. That sounds evasive. It is not. Those three variables move the number by an order of magnitude, and anyone who quotes you a total before knowing them is guessing.
What is not evasive is the structure. eDiscovery spend is predictable in shape even when it is unpredictable in size. Every matter distributes cost across the same handful of phases, in roughly the same proportions, and the phase that dominates is almost always the same one. Understand that structure and you can build a budget that survives contact with a real matter. You can also tell the difference between a quote that is expensive and a quote that is badly scoped.
The short version
For a US commercial matter in the mid-2020s, typical market ranges look roughly like this. Every figure varies by vendor, region, data type, and how hard you negotiate. Treat them as orientation, not as a rate card.
- Processing: roughly $20 to $50 per gigabyte ingested, and falling. Some providers now bundle it, or price it near zero to win the hosting.
- Hosting: roughly $5 to $25 per gigabyte per month of hosted data, often with per-user platform fees on top. This is the line that accrues quietly for the life of the matter.
- Forensic collection: commonly $200 to $400 per hour for a qualified examiner, or a fixed fee in the $300 to $1,000 range per custodian for standard cloud sources.
- Managed document review: $30 to $60 per hour for contract reviewers, or $0.50 to $2.00 per document where a provider prices by volume rather than by time.
- Project management: $150 to $300 per hour, and you will use more of it than you think.
Put those together and the totals spread widely. A single-custodian employment matter with a targeted collection can close out in the low five figures. A twenty-five custodian commercial dispute with a full linear review lands in the high six figures without anyone doing anything wrong. The difference between those two outcomes is rarely the rate card. It is the volume.
Where the money goes, phase by phase
The EDRM model is the standard way to describe the discovery lifecycle. It also works as a budget skeleton.
Identification and preservation
Usually the cheapest phase in direct spend, and the most expensive one to get wrong. Issuing holds, interviewing custodians, and suspending auto-deletion consume counsel and IT hours rather than vendor dollars. Organizations with a current data map and a hold process that already reaches IT spend a fraction of what improvising teams spend, because they are not re-deriving the custodian list from scratch every time. Budget this as hours, not as an invoice line.
Collection
Priced two ways. Forensic imaging of laptops, phones, and servers is billed hourly by the examiner, plus travel where a device cannot be shipped. Remote collection from Microsoft 365, Google Workspace, and Slack is increasingly priced per custodian, because much of the work is scripted. The cost driver here is not technical difficulty. It is the number of separate systems in scope. Each new source adds setup, validation, and chain-of-custody documentation whether it yields ten documents or ten thousand.
Processing
Per gigabyte, almost always. Processing is where collected data is expanded, de-duplicated, indexed, and made searchable. It is the phase where volume first becomes visible, and where the number on the invoice diverges from the number in your head. A 500 GB collection may reduce to 120 GB of reviewable data after de-NISTing, de-duplication, and date filtering. Find out which of those numbers a quote is built on.
Hosting
Per gigabyte per month, which makes it the only cost that grows while nothing happens. A matter hosting 200 GB at $10 per GB per month costs $2,000 a month, or $72,000 across a three-year case. Nobody approves $72,000. Everybody approves $2,000 a month, thirty-six times. Ask for the projected total over the expected life of the matter, and ask what the rate does when the case goes quiet.
Review
The dominant line item, and the subject of the next section.
Production
Comparatively small: image conversion, endorsement, load-file generation, and the mechanics of getting the set out the door. It gets expensive in exactly one situation, which is when the format was never agreed and the set has to be produced twice. That is a strong argument for settling formats and metadata fields at the Rule 26(f) conference, rather than discovering the disagreement after the first production has already gone out.
Processing and hosting are priced per gigabyte, so they are what people benchmark. Review is priced per hour or per document, and it is what actually determines the budget. Negotiating your per-GB rate while leaving review volume untouched is arguing over the wrong number.
Why review dominates
Arithmetic, mostly. Take a 100 GB review population. Depending on file types, a gigabyte of business data commonly yields somewhere between three thousand and ten thousand documents, so call it 500,000 documents. A competent contract reviewer moves through perhaps 40 to 60 documents an hour on substantive first-pass review. At 50 an hour, that is 10,000 review hours. At $40 an hour, it is $400,000 in review labor before a single privilege call has been quality-checked.
Now compare the other lines on the same 100 GB. Hosting at $10 per GB per month is $1,000 a month. Processing was a one-time charge in the low five figures. Review is not one line among several. On a linear review it routinely accounts for half to three-quarters of total discovery spend, which means nearly every decision in the matter is really a decision about how many documents reach a reviewer.
That is why associate review at $400 an hour is so rarely the right instrument for first-pass work, and why the choice between linear review and technology-assisted review is the largest single budget lever available to you.
Three pricing models, and when you will see each
- Per gigabyte. The default for processing and hosting. Easy to quote and easy to compare, which is exactly why vendors like it and why it can mislead. The same 100 GB produces wildly different document counts depending on whether it is Outlook archives or scanned PDFs. Always ask whether the rate applies to data ingested or to data promoted into review.
- Per custodian. Common for remote cloud collection, for early data assessment, and increasingly for all-in bundles quoted per custodian per month. It is genuinely useful for budgeting, because custodian counts are knowable early and gigabytes are not. Read the assumed data ceiling per custodian carefully; the bundle usually stops being a bundle above it.
- Per document. Used in managed review and some all-in offerings. It transfers reviewer-speed risk to the provider, which is worth paying for. It also makes culling urgent, because under per-document pricing every document you failed to remove is a line on the invoice.
Hourly pricing sits alongside all three for forensic work, project management, and expert time. A quote that mixes models is not a red flag. A quote that will not tell you which model applies to which phase is.
Volume assumptions drive everything
Every figure above gets multiplied by an assumption about how much data exists. That assumption is usually made early, casually, by someone with incomplete information, and it is rarely revisited once it is in the spreadsheet.
The corrective is early data assessment: pulling volume statistics, date distributions, file-type breakdowns, and custodian-by-custodian counts before committing to a collection scope. It costs something, typically a few thousand dollars on a mid-sized matter, and it routinely changes the plan. Learning that two of your twelve custodians hold 60 percent of the data, or that a third of the collection predates the relevant period, is worth considerably more than it costs to find out.
The levers that actually reduce eDiscovery cost
- Targeted collection. Collect the custodians and date ranges the matter needs, on a documented rationale, rather than everything within reach. Over-collection is the original sin: it inflates processing, hosting, and review at the same time.
- De-duplication and email threading. Global de-duplication across custodians plus thread suppression routinely removes 30 to 50 percent of a review population before anyone reads anything. This is housekeeping, not a negotiation with the other side.
- De-NISTing and file-type filtering. System files, application binaries, and interface graphics are not evidence. Removing them is standard practice and belongs in every processing specification.
- Search-term testing. Test proposed terms against the actual corpus before agreeing to them. One unqualified term, such as the company's own name, can add hundreds of thousands of documents. Once it is in a stipulated protocol, you own it.
- TAR and continuous active learning. On populations above roughly 50,000 documents, prioritized review under a validated protocol typically reduces the number of documents a human reviews by a large multiple. The savings are real. The defensibility depends entirely on documenting the protocol and the validation.
- Phased scope. Agreeing to review a first tranche of custodians and revisit scope afterwards stops you paying for the tail before anyone knows whether the tail matters.
The costs nobody puts in the budget
- Project management. Real hours on every matter, spent on scoping calls, exception handling, load-file troubleshooting, and keeping counsel and vendor pointed the same direction. A provider who quotes no project management either buried it or will not deliver it.
- Re-collection. The most expensive thing in eDiscovery is doing the collection twice. It happens when a source was missed, when the first pass was not forensically sound, or when the custodian list was wrong. You pay the entire downstream stack again, under deadline, at rush rates.
- Expert testimony. If your process is challenged, someone credible has to explain it. Testifying experts commonly bill in the $400 to $900 per hour range, and preparation runs far longer than the testimony.
- Privilege logging. Consistently underestimated, especially where a document-by-document log is required instead of a categorical one.
- Sanctions exposure. Not a budget line, but often the largest number in the analysis. A Rule 37(e) dispute costs more in motion practice alone than the preservation step that would have prevented it.
- Data left hosted. Storage that keeps billing after a matter closes because no one made the decision to take it down.
Cost is a legal argument, not just a budget
Rule 26(b)(1) of the Federal Rules of Civil Procedure makes proportionality part of the definition of what is discoverable at all. Discovery must be proportional to the needs of the case, considering the amount in controversy, the parties' resources, the importance of the issues, the importance of the discovery in resolving them, and whether the burden or expense outweighs the likely benefit. Cost is not a complaint you make to a court. It is an element of the standard.
Which means your cost model is evidence. "This request is burdensome" persuades nobody. "This request reaches four systems and roughly 640,000 documents, at a review cost we have modeled and can show you, in a case with this much in controversy" is an argument a court can act on. It is only available to a party that did the volume work early.
The same logic runs through cost-shifting. In Zubulake v. UBS Warburg LLC, 217 F.R.D. 309 (S.D.N.Y. 2003), the court replaced the earlier eight-factor test from Rowe Entertainment, Inc. v. William Morris Agency, Inc., 205 F.R.D. 421 (S.D.N.Y. 2002) with a seven-factor test for deciding when the cost of producing inaccessible data should shift to the requesting party. It also weighted the factors in descending order of importance, treating the first two, which it called the marginal utility test, as the most important. When the same court applied that test later in the litigation, it shifted 25 percent of the cost of restoring backup tapes to the plaintiff and left every other cost with the producing party. See Zubulake v. UBS Warburg LLC, 216 F.R.D. 280 (S.D.N.Y. 2003).
Two lessons survive from that. Cost-shifting is available, and it is partial. A producing party should expect to carry most of its own discovery cost, which is exactly why controlling volume beats litigating over who pays.
Not "what is your per-GB rate." Ask instead: on a matter like this one, what do you expect the total to be, what volume assumption is that built on, and what happens to the number if the assumption is off by half? A provider who can answer has scoped matters like yours before. A provider who cannot is selling you a price list.
Building a number you can defend
A usable eDiscovery budget has four inputs and one honest caveat. The inputs are custodian count, data volume per custodian, expected reduction from culling, and the review method. The caveat is that the first two are estimates until early data assessment turns them into measurements.
Build the forecast as a range with the assumptions written beside it, not as a single figure with false confidence. Re-forecast after processing, when you know the real document count rather than the guessed one. Then keep the model. The second matter is far cheaper to scope than the first, provided the playbook records what the last one actually cost.