WEBVTT

00:00:00.001 --> 00:00:07.460
Welcome to Legal Prompting, the podcast dedicated to legal methods in the age of artificial intelligence.

00:00:07.900 --> 00:00:09.739
I'm Nicola Fabiano.

00:00:10.119 --> 00:00:15.420
In the previous episode, we talked about RUG and its risks in the legal domain.

00:00:15.800 --> 00:00:17.979
Today, we take a step forward.

00:00:18.500 --> 00:00:23.399
We discuss two techniques that change how we interact with models,

00:00:24.440 --> 00:00:26.940
chain of thought and few-shot prompting.

00:00:26.979 --> 00:00:28.860
Let's start with the premise.

00:00:28.860 --> 00:00:31.920
A model does not reason like a lawyer.

00:00:32.619 --> 00:00:36.159
It produces plausible answers, not reasoned conclusions.

00:00:37.180 --> 00:00:40.779
But we can steer its process with precise techniques.

00:00:41.619 --> 00:00:43.340
Let's look at them one at a time.

00:00:43.819 --> 00:00:50.000
Chain of thought is the explicit request to lay out the logical steps before the conclusion.

00:00:51.060 --> 00:00:55.020
Instead of simply asking, is this clause valid, we ask,

00:00:55.020 --> 00:00:57.959
examine the clause step by step,

00:00:59.380 --> 00:01:00.939
identify the applicable provision,

00:01:01.840 --> 00:01:03.000
verify the conditions,

00:01:03.819 --> 00:01:04.779
assess the exceptions,

00:01:05.400 --> 00:01:07.819
only then formulate the conclusion.

00:01:08.220 --> 00:01:10.319
Let's take a concrete example.

00:01:10.860 --> 00:01:14.459
We have a clause on data transfers outside the EU.

00:01:15.720 --> 00:01:19.720
Instead of a blunt question, we structure the reasoning as follows.

00:01:19.720 --> 00:01:26.160
First, identify the legal basis for the transfer under Chapter 5 of the GDPR.

00:01:27.519 --> 00:01:29.879
Second, verify the safeguards adopted.

00:01:30.839 --> 00:01:33.480
Third, consider the SHREMS 2 case law.

00:01:34.819 --> 00:01:36.800
Fourth, conclude on compatibility.

00:01:37.279 --> 00:01:39.300
The result changes right away.

00:01:39.639 --> 00:01:42.000
The model makes the path explicit.

00:01:42.800 --> 00:01:47.720
And this allows us to verify where the reasoning holds and where it slips.

00:01:47.720 --> 00:01:52.419
It's the difference between a statement and an argument.

00:01:52.879 --> 00:01:53.980
A warning, though.

00:01:54.519 --> 00:01:57.879
Chain of thought does not guarantee that the reasoning is correct.

00:01:58.419 --> 00:02:01.260
It only guarantees that it is made explicit.

00:02:02.059 --> 00:02:06.040
The model can build coherent logical steps on flawed premises.

00:02:06.739 --> 00:02:09.399
Our verification remains indispensable.

00:02:09.979 --> 00:02:12.279
Let's move to few-shot prompting.

00:02:13.279 --> 00:02:17.240
Here, we give the model examples of how we want it to answer

00:02:17.240 --> 00:02:19.720
before posing the actual question.

00:02:20.639 --> 00:02:24.580
Two, three, at most five well-chosen examples.

00:02:24.960 --> 00:02:28.100
In the legal context, it works like this.

00:02:28.679 --> 00:02:33.979
Do we want the model to analyze a decision of the supervisory authority

00:02:33.979 --> 00:02:36.800
in accordance with a precise structure?

00:02:37.720 --> 00:02:43.279
We show two or three analyses already done with that structure.

00:02:43.820 --> 00:02:45.740
Then we submit the new decision.

00:02:46.460 --> 00:02:49.240
The model tends to replicate the format.

00:02:49.740 --> 00:02:52.820
The quality of the examples is everything.

00:02:53.679 --> 00:02:56.500
Generic examples produce generic results.

00:02:57.440 --> 00:03:00.460
Precise examples with the correct technical language

00:03:00.460 --> 00:03:03.240
and the argumentative structure we need

00:03:03.240 --> 00:03:06.679
produce much better aligned outputs.

00:03:07.600 --> 00:03:11.839
It's instruction by demonstration, not by explanation.

00:03:12.419 --> 00:03:14.479
The two techniques can be combined.

00:03:15.279 --> 00:03:18.699
We can provide examples of chain of thought reasoning

00:03:18.699 --> 00:03:22.100
already structured in a few-shot mode.

00:03:22.419 --> 00:03:25.520
The model learns both the format and the method.

00:03:26.440 --> 00:03:31.279
For analyzing complex decisions, it's often the most effective combination.

00:03:31.279 --> 00:03:33.440
Let's move to the limits.

00:03:34.160 --> 00:03:36.800
The first is the length of the context.

00:03:37.779 --> 00:03:39.759
Every example takes space.

00:03:40.779 --> 00:03:44.759
In long documents, we have to balance the number of examples

00:03:44.759 --> 00:03:47.919
and the complexity of the text we analyze.

00:03:48.300 --> 00:03:50.960
The second limit concerns bias.

00:03:51.460 --> 00:03:54.820
If all our examples follow a certain interpretation,

00:03:55.479 --> 00:03:59.479
the model will apply it even where it's not appropriate.

00:03:59.479 --> 00:04:03.080
Examples shape reasoning, not just form.

00:04:04.000 --> 00:04:08.080
Let's choose representative examples, not convenient ones.

00:04:08.619 --> 00:04:10.979
The third limit is the most insidious.

00:04:12.000 --> 00:04:14.820
Unexplicit reasoning looks more reliable,

00:04:15.559 --> 00:04:18.679
but plausibility is not legal correctness.

00:04:19.100 --> 00:04:22.600
A well-built argument on a non-existent rule

00:04:22.600 --> 00:04:26.779
remains a hallucination, only more convincing.

00:04:26.940 --> 00:04:28.260
A practical caution.

00:04:28.260 --> 00:04:32.899
When we use these techniques for real legal work,

00:04:33.160 --> 00:04:36.700
we document everything, the prompt, the examples,

00:04:37.000 --> 00:04:39.420
the output, and our verification.

00:04:40.279 --> 00:04:43.279
It is part of the governance of AI use,

00:04:43.480 --> 00:04:47.100
and it will be increasingly relevant with the AI act.

00:04:47.179 --> 00:04:48.600
One last thought.

00:04:49.720 --> 00:04:52.760
Chain of thought and few-shot are not tricks.

00:04:53.739 --> 00:04:56.660
They are the way we translate our legal method

00:04:56.660 --> 00:04:59.619
into instructions understandable to the model.

00:05:00.600 --> 00:05:03.899
The clearer our method, the better the techniques work.

00:05:04.260 --> 00:05:07.700
In the next episode, we will apply these techniques

00:05:07.700 --> 00:05:10.579
to the analysis of contracts and clauses.

00:05:11.179 --> 00:05:15.140
We will see how to build effective checklists,

00:05:15.940 --> 00:05:17.660
how to compare versions,

00:05:17.739 --> 00:05:21.820
and which limits remain beyond the model's reach.

00:05:22.160 --> 00:05:23.399
Thanks for listening.

00:05:23.399 --> 00:05:26.399
If you find this podcast useful,

00:05:27.239 --> 00:05:30.500
share it with colleagues interested in legal methods

00:05:30.500 --> 00:05:32.500
in the age of AI.

00:05:33.399 --> 00:05:37.299
We'll meet again in the next episode of Legal Prompting.

