CASUS
Open navigation

Product

Expand Product

Customers

Open Customers menu

Company

Open Company menu

Security

Pricing

Casus Logo

CASUS Blog

In-House Legal AI: A 6-Week Pilot Playbook

Last updated on

by

Fabian Staub

Fabian Staub

|

Co-Founder & CEO

In-house Legal AI should produce a decision through the pilot, not a collection of impressive examples. In six weeks, a legal team can define one or two workflows, approve their data and review rules, test them on a fixed set of real tasks and decide whether each use case should scale, change or stop.

The six artefacts the pilot must produce

Before the calendar starts, name the outputs. A useful pilot leaves behind:

  1. a use-case matrix;

  2. a data and risk classification;

  3. a fixed evaluation set;

  4. review and escalation rules;

  5. a measurement sheet;

  6. a written go/no-go decision.

These artefacts keep the exercise focused. They also make the outcome auditable when stakeholders remember the most successful demo but not the failures or review effort.

The broader guide to AI in legal departments covers governance and scaling. This playbook concentrates on execution over six weeks.

Week 1: define the use case

List recurring tasks and score them for frequency, structure, data sensitivity, reviewability and potential operational value. Do not choose only by perceived time consumption. A task can take many hours yet remain unsuitable because inputs vary too widely or quality is difficult to judge.

Select no more than two workflows for the first pilot. Describe each as an input, transformation, expected output and responsible reviewer. Examples might be a first pass over recurring agreements or extraction of defined terms from a controlled document set.

Write explicit exclusions. The pilot should state which matters, data classes and output types are not included. This is easier to enforce than a general instruction to “use good judgement”.

Use-case matrix

For each candidate, capture:

Field

Example of the decision required

Input

defined agreement type or approved public sources

Output

issue list, structured table or sourced note

Frequency

enough tasks during the pilot to observe a pattern

Reviewability

clear standard and identifiable reviewer

Data

permitted classification and required safeguards

Value

effort, cycle time, consistency or visibility

The selected use case should have an owner who can make operational decisions, not only sponsor the project.

Week 1: complete the governance gate

Map the data path from source system to tool and back. Identify what is uploaded, which metadata is created, where output is stored and who can access it. Review current provider documentation and contracts for the intended workflow.

Set practical rules: approved accounts, permitted data, prohibited data, retention expectations, required access controls, incident channel and the event that suspends the pilot. A safe failure mode is part of the design.

Start with synthetic, public or appropriately prepared material while the review is incomplete. The pilot should not use confidential data merely because a test account is technically available.

Week 2: build the evaluation set and baseline

Choose completed examples before the tool is used. The expected output and known difficult points should already be available. Include normal cases, an edge case and at least one example that tests whether the system admits uncertainty.

Record the existing workflow with consistent boundaries. Include preparation, active work, review, correction and transfer into the final system. If the old process and new process are timed differently, the comparison will be misleading.

Five to ten tasks per use case may produce a useful first signal when the work is recurring and comparable. It does not support a market-wide percentage. Report the sample and variation honestly.

Weeks 2–3: define review rules and train the test group

Create a one-page workflow card. It should show the approved input, expected output, mandatory checks, prohibited shortcuts and escalation route. Training uses the same tasks and systems that participants will use during the pilot.

Demonstrate failure detection. Participants should practise finding unsupported statements, missed clauses, wrong source context and formatting that obscures uncertainty. A training session focused only on successful prompts gives the team no method for safe review.

Assign primary and backup reviewers. If no one has time to review results, the use case is not operationally viable even if the technology works.

Weeks 3–5: run controlled live tasks

Use the fixed workflow on real tasks within the approved scope. Record every run, including abandoned or unusable results. Avoid changing the task definition after each failure; otherwise no comparable pattern can emerge.

For each run capture:

  • task category and complexity;

  • total active time;

  • review and correction time;

  • material errors or omissions;

  • whether the output met the quality threshold;

  • whether it was used, revised or discarded;

  • user notes about friction and handoffs.

Short weekly reviews can resolve unclear instructions and operating problems. Keep changes visible. If the workflow is materially redesigned, mark a new test phase instead of mixing the results.

Weeks 4–5: examine the failures

Failures are the most valuable evidence in a small pilot. Classify them rather than treating them as anecdotes. Was the input incomplete, the task poorly specified, a source unavailable, the answer unsupported, the review rule unclear or the product workflow cumbersome?

Distinguish a correctable operating issue from a structural mismatch. Training may solve inconsistent instructions. It will not create missing source access or remove a data-processing concern. The decision log should show which response is appropriate.

Repeated material errors should trigger the predefined suspension rule. The pilot exists to learn within a controlled boundary, not to defend continuation.

Weeks 5–6: evaluate the complete workflow

Summarise adoption, quality, effort and outcome separately. Do not lead with the number of logins or generated outputs. Show the share of tasks meeting the quality threshold, the review burden, the range of total effort and the observed operational effect.

The measurement guide for AI time savings in legal work explains why review and rework must remain in the calculation. If the sample is small or tasks differ substantially, report that limitation.

Ask test users and reviewers different questions. Users can identify friction and learning needs; reviewers see error patterns and quality risk. Internal clients may reveal whether cycle time or clarity actually improved.

End of week 6: make the decision

Use predefined outcomes for each workflow:

  • Go: quality and operating criteria are met; expand with named conditions.

  • Adjust: the use case is promising but input, scope, training or controls must change before another test.

  • Stop: the workflow does not meet the quality, effort or governance threshold.

The decision document states the scope tested, sample, evidence, unresolved risks, owner and next review date. A conditional go should list the conditions explicitly. Do not convert a successful pilot for one task into approval for every Legal AI use case.

Prepare the scaled operating model

If the decision is go, define access provisioning, onboarding, support, monitoring and change review before adding users. Decide which product or provider changes require the workflow to be reassessed.

Maintain a small set of reference tasks for regression checks. When the system or working method changes, rerun them. This turns the pilot from a one-off procurement event into an operating discipline.

CASUS can be evaluated for document review, legal research and related workflows. Teams can start with their own controlled task set, provided the data and review gates are already defined.

Frequently asked questions

Why use a 6-week pilot?

It creates enough structure for governance, training, live tasks and a decision while preventing an open-ended experiment. Task volume and evidence quality remain more important than the exact duration.

How many use cases should the first pilot include?

One or two. A narrow scope makes it possible to define inputs, review standards and measurements precisely.

Who should participate?

Include the legal use-case owner, representative users, named reviewers and the functions needed for data, security, procurement and technical operation.

What is a material error?

Define it before testing. It may be an unsupported legal proposition, a missed high-priority clause, incorrect source context or another defect that makes the output unusable without substantive correction.

What if the pilot saves time but quality declines?

The workflow should not scale unless it meets the agreed quality threshold. Speed is an outcome only after acceptable quality and control are established.

Your Legal AI Associate.

Supported by Innosuisse, the Swiss Innovation Agency
Capterra rating: 5 out of 5
Spin-off from the University of St. Gallen

CASUS Technologies AG Beethovenstrasse 48
8002 Zurich
Switzerland
contact@getcasus.com

Ask your favorite AI about CASUS

ChatGPT
Claude
Perplexity

Your Legal AI Associate.

Supported by Innosuisse, the Swiss Innovation Agency
Capterra rating: 5 out of 5
Spin-off from the University of St. Gallen

CASUS Technologies AG Beethovenstrasse 48
8002 Zurich
Switzerland
contact@getcasus.com

Ask your favorite AI about CASUS

ChatGPT
Claude
Perplexity

Your Legal AI Associate.

Supported by Innosuisse, the Swiss Innovation Agency
Capterra rating: 5 out of 5
Spin-off from the University of St. Gallen

CASUS Technologies AG Beethovenstrasse 48
8002 Zurich
Switzerland
contact@getcasus.com

Ask your favorite AI about CASUS

ChatGPT
Claude
Perplexity