DOES AI REALLY LOWER THE COST OF DOCUMENT REVIEW?
NOT ALWAYS.
Well, as always in legal issues, the answer is “it depends.” While recent cases like Schulte v. LinkedIn Corp. signal growing court acceptance of generative AI under existing technology-assisted review frameworks (cf. https://www.jdsupra.com/legalnews/court-greenlights-generative-ai-for-6075451/), the financial benefits are still suspect, particularly for small law firms.
Why? Because document richness and corpus size dictate cost-effectiveness more than any other factors. In fact, AI can actually increase expenses for small-volume matters or through rigid per-seat pricing products.
How does document richness affect the cost of AI review?
Document richness, defined as the percentage of responsive documents within a dataset, affects the cost-effectiveness of AI review. The impact varies depending on whether the richness is high or low:
High Richness
High richness impacts ROI because AI and machine learning methods provide the greatest return on investment in matters with a high percentage of responsive documents. This is because improvements in precision, a key strength of advanced AI, are most valuable when there is a large volume of relevant material to identify and classify.
Low Richness
Low richness, however, shows a diminishing advantage as the overall review-cost component shrinks. This mutes the advantage of high precision and allows “cheaper-to-run” methods to become cost-competitive compared to more advanced AI tools.
Low richness can actually make Generative AI (GenAI) more expensive than traditional Technology-Assisted Review (TAR) or manual review. GenAI typically reviews every document individually, often costing between $0.11 and $0.50 per document.
This cost scales with the total number of documents regardless of how many are actually relevant. In a collection with very few responsive documents, the expense of running an LLM inference call on every file may exceed the labor savings.
Ultimately, for small firms or smaller matters, AI review, especially GenAI, may only lower costs if the matter is large enough and the workflow focuses on simple responsiveness rather than complex privilege reviews.
IS THERE A SOLUTION?
Yes, and it’s called Early Data Assessment (EDA). It helps manage document richness by performing volume-estimating and culling before the primary review begins.
How does that occur?
Several ways, including:
Increasing the Percentage of Responsive Docs
Document richness is defined as the percentage of responsive documents within a corpus. By culling irrelevant data early, EDA effectively increases the richness of the remaining dataset.
Maximizing AI Return on Investment (ROI)
Improvements in precision for AI methods provide the greatest return on matters with higher richness. By using EDA to remove “noise” and increase the concentration of responsive documents, firms can better leverage the high-precision capabilities of advanced AI tools.
Avoiding “Cost Traps”
Without EDA, a dataset may have low richness, which reduces the precision advantage and allows cheaper, less sophisticated methods to catch up in cost-effectiveness.
In cases of low richness, GenAI can become particularly expensive because it typically reviews every document individually at the previously mentioned price range of $0.11–$0.50 per document.
Culling via EDA is essential to prevent paying for individual AI analysis on non-responsive files.
Enabling Proportionality Analysis
EDA allows for volume estimation, which can materially influence proportionality decisions in court, as seen in cases like Schulte v. LinkedIn Corp.
In summary, EDA acts as a filter that refines the corpus, ensuring that the subsequent AI review is focused on a “richer” set of data, thereby making the use of advanced technology more economically viable.
SO, WHAT ARE THE BEST PRACTICES FOR SMALL-FIRM BUDGETING FOR AI AND DOCUMENT REVIEW?
The key is to avoid the AI “cost traps” mentioned above, where technology expenses can actually exceed labor savings.
Here are some steps I recommend:
1. Prioritize Usage-Based Pricing
Small firms should generally avoid per-seat subscription models, which often range from $200 to $500 per lawyer per month. This structure scales poorly for firms with 1–10 lawyers who may have inconsistent usage.
Instead, budget for usage-based models that price out at $0.02–$0.05 per document, as these align costs directly with the workload of specific matters.
2. Apply Volume Thresholds for AI Use
Budgeting for AI is most effective when applied to the right size of matter:
Small Matters
Avoid expensive GenAI workflows in cases with less than 10,000 documents, where fixed costs often outweigh labor savings, leading to negative ROI.
Medium-Volume Matters
Cases with 10,000–50,000 documents are the “sweet spot” where AI can eliminate days of manual review and deliver major savings.
3. Invest in Early Data Assessment (EDA)
Budgeting for EDA is essential before committing to a full AI review. By removing irrelevant data, you are not paying for AI to analyze “noise.”
Furthermore, increasing document richness through EDA maximizes the return on investment for high-precision AI tools.
4. Target High-ROI Workflows
Small firms see the strongest ROI when budgeting AI for specific high-billable-rate tasks such as:
Contract Review and Due Diligence
Replacing a $250–$400/hr lawyer with a low-cost inference call is a clear economic win.
Simple Responsiveness
AI is more cost-effective for simple responsiveness calls than for complex, privilege-heavy matters that still require significant human validation.
5. Account for Hidden Operational Costs
A comprehensive budget must look beyond vendor fees and include hidden operational costs, such as:
- Training and Onboarding: Time spent learning new platforms.
- Workflow Redesign: Adjusting firm processes to accommodate AI.
- Peer Review/Validation: Human oversight required to ensure a “reasonable” production, as seen in the Schulte v. LinkedIn Corp. case.
IN SUMMARY
Achieving a positive ROI requires matching the specific pricing structure and workflow to the scale of the legal matter.
Prioritize Early Data Assessment and usage-based models to ensure that technological adoption results in reasonable and proportional discovery outcomes.





