Harvey and Claude get far more useful when you shape the context you give them, not when you chase a smarter model.
Ashkaan Hassan · We Solve Problems · prepared for RAMO
I'm CEO of We Solve Problems, a managed IT firm in Los Angeles, and I run an AI operations practice for public companies. I'm also an attorney and a chief compliance officer, so I come at this as a governance problem first and a speed problem second.
of automations shipped, mine and my clients'
running unattended, nobody watching
systems wired together, ticket queue to bank
running IT for regulated businesses
Not chat prompts. Autonomous agents. One of them is a chief of staff I call Lucius: I hand it a whole task, it makes the calls I would have made, and it reports back with the log.
It runs my company and my house. It also built this deck. None of it is a demo.
You open Harvey, or Claude.
You explain your work from scratch.
You get a decent answer.
You close the tab.
Repeat tomorrow.
Sound familiar?
An AI without your context is a brilliant lateral hire with amnesia every morning.
You do all the work bringing it up to speed. Every single time.
The model is already smart enough. The problem is that it doesn't know you.
What your team could take on was always bounded by who you could hire. And nobody you hired was useful on day one. They got useful because somebody trained them.
You can buy that capability by the unit now. It still shows up knowing nothing about how you work.
The bottleneck moved.
“Can I afford to hire for this?”
“Have I trained it the way I'd train a person?”
So the skill isn't prompting. It's onboarding.
Hands up for yours. No wrong answer — I just need to know where the room starts.
You ask, you get an answer, you close the tab. Tomorrow starts from scratch.
You built something durable, and next week you pick it up where you left off.
You handed it to other people, and it works for them without you in the room.
It is how the company does that job now. It has an owner and a budget.
Wherever your hand went up, you leave today one rung higher.
Level-setting, so we spend the next hour arguing about the same things.
The unit AI reads and writes in — roughly three-quarters of a word. Every size limit and most of the bill is counted in tokens.
Everything the model can see when it answers: your question, plus whatever has been handed to it. It is finite, and it is the whole ballgame.
Pointing it at your real documents so the answer comes from your matters instead of the open internet. In Harvey that layer is Vault and Knowledge — your files, plus the legal sources it is licensed to read.
How far it goes before it checks with you. Suggests → acts with your approval → acts and tells you after. A dial you set, not a property of the tool.
A model given instructions, tools, and permission to act — not just answer. Turn that dial past “suggests” and what you have is an agent.
Not three products to choose between. Three points on one dial — how much of the work you hand over.
You ask, it answers, you close the tab.
You get an answer.
You describe an outcome. It works in the cloud with your laptop shut and comes back when it’s done.
You get a document.
You hand over the task. Several agents split it, work in parallel against your own documents, and hand back work product with the citations attached.
You get something to review.
Same model, same failure, same fix at every point on the dial: each is only as good as the context you hand it.
Let's put the uncomfortable one first, because every firm asks it and nobody asks it out loud.
It is a fair question and the honest answer is yes — for that task. The four hours of first-draft time were never the thing the client valued anyway.
Which is why the firms that stall are the ones treating this as a discount, instead of asking what the freed hour is worth.
The hour does not disappear. It moves — out of drafting and document review, into negotiation, structuring and judgment. That is the work the client is actually paying for.
The firms I watch take on more matters with the same people, and postpone the hire they thought they needed this year.
Nobody here is being asked to bill less. You are being asked what you would do with the hour back.
Four questions. None of them is a question about the model — every one is a question about the account you opened and the rules this firm set.
The firm's licensed Harvey and the free account on your phone are not the same product with a different logo. Same model, completely different agreement.
Whether it is retained, and whether it trains anything. The answer lives in the contract the firm signed, not in the marketing page.
A shared vault is a permissions decision, not a filing decision. Matters that are walled off from each other on the server have to stay walled off here.
Harvey returns citations for a reason. An answer you cannot walk back to a document is not something you put your name on.
Get these four right once, as a firm, and nobody has to relitigate them per matter. Get them wrong once and it is a privilege problem, not an IT problem.
And it is visible: Harvey's Command Center shows the firm what is being used, by whom, and where it is actually working.
The difference between an AI that disappoints you and one you would not give up is not a smarter model. It is how much of your world it has been told.
Shape the context and it stops answering like a stranger. It starts answering like someone who works here.
And you control that entirely. There are exactly six ways to shape it, and the rest of today is those six.
Six ways to hand it your context. Everything from here is one of these, and they work the same way in Harvey and in Claude.
How it behaves every time. Your voice, your standards, your limits.
Your precedents, your matters, your house positions. Harvey ships this as Vault and Knowledge.
How far it goes before checking with you — and whether it waits to be asked at all.
The systems it can actually reach into and act on, not just read.
Decisions that survive the session, instead of scrolling away.
A correction becomes a rule. The lever that keeps the other five current.
The four rungs from the show of hands are where you are. These six are how you climb. Pull one and you feel it; pull all six and what you built is good enough to hand to someone else.
Six levers that shape context. They work the same way in Harvey and in Claude.
It behaves the same way every time — your way. How you communicate, what to reach for, what's off limits.
In Harvey this is a playbook — your positions, your fallbacks, your redlines, versioned so you can see what changed. In Claude it's a project's instructions. Same lever, two homes.
Write it once. Every request after that carries it.
Four things, and they're all things you already say out loud to a first-year.
Our voice, our defined terms, our level of hedge. This is the line that stops you rewriting every draft so it sounds like this firm instead of a form book.
Which document is the truth for a given question. The current form lives in the precedent bank, not in whichever deal you happened to have open. Name the source and it stops guessing.
Your house process. Redline against our form, not theirs. Check the defined term against the executed version, not the draft. Cite the clause you relied on.
The hard stops. Nothing goes to a client, nothing gets filed, nothing gets sent to opposing counsel without a lawyer saying yes. Write the limits down or you'll be enforcing them by hand.
Different agents can carry their own playbook. The one doing clearance research gets different rules than the one near a signature page.
The demo everyone shows is the one that produces something. In a law firm, the demo that matters is the one that refuses.
And a refusal like this is not the model being careful. It is a line somebody wrote in plain English — so it holds the same way every time, for everyone, whether or not anyone is watching.
Which means the interesting test is trying to talk it out of the rule. That is ten minutes of your own time, and it is the ten minutes worth spending.
“Answer the studio’s comments on the deal and send it back to them.”
I can draft it. I can’t send it. That goes to the other side, so it needs a lawyer on it first. What I have ready is the response against our playbook, with the two terms we don’t concede flagged, sitting in the matter for you.
It always has exactly the right rules for what it's touching — without you asking for it.
You don't want one giant instruction sheet. You want one playbook per kind of work — because the rules for a distribution deal are not the rules for a music clearance.
It never loads everything at once. It stays fast and focused by only pulling what the moment requires.
It knows your world — your clients, your precedents, your house positions. Harvey ships this as Vault, which holds your documents in one searchable place, and Knowledge, the legal sources it is licensed to read. It's why the draft comes back sounding like this firm instead of a form book.
Name a client — it has the file. Ask about a term — it has the last four times you negotiated it. Context shows up automatically.
The dial from earlier, turned up — and three words that get used interchangeably once you do. The agent is the worker. That worker on a schedule is an automation. How far it goes before checking with you is autonomy, a dial you set.
Work happens while you sleep. In Harvey that's an agent: you hand over the task, several of them split it and work in parallel, and you get back work product with the citations on it. Harvey's own words for where the dial sits are the right ones — shape the plan before it runs, review the work before it goes out.
I have hundreds of these running right now, mine and my clients'.
The work doesn't live in one place. It lives here. An agent is only useful to the degree it can reach into these and act, not just read about them.
Harvey already comes to you inside Word and Outlook. Some of the rest you wire — which is exactly what the Clio move is for. Either way, that wiring is the ceiling on what you can delegate.
“Why did we give on the holdback that time?” Two years and one associate departure later, it still knows. That only works if the answer was written down on the day it was decided.
Every initiative has one home: where it stands, what's open, what's next. The agent updates it as the work moves, so nobody rebuilds status from memory.
Every session ends with what was decided, what changed, and what broke. Searchable a year later, when somebody asks why it was built this way.
Vault already holds the documents where your decisions actually landed, and Claude keeps what you told it from one session to the next. Neither is a filing system anyone has to maintain by hand.
The memory gets better over time, not worse. That's the opposite of how it works with people.
The one lever that keeps the other five current. A correction becomes a rule, so your context updates itself. Three parts — miss any one and it's dead.
Did it find out how it went?
The correction you gave. The deal that closed. The alert that fired.
Did the lesson get written down?
A rule in a file. Not a chat message that scrolls away and is gone by morning.
Does the next run read it?
The rule loads before the work starts. So the mistake can't happen twice.
Signal without capture is a lesson you forget. Capture without recall is a file nobody reads. You need all three.
The loop isn't a diagram. This is one correction I gave once, still holding today.
Once, it went to put a change live on its own. I stopped it: nothing goes live until I personally say so.
It reads that line before every session. It has asked me first every time since — and I never had to repeat myself.
I didn't program that. I corrected it once, and the system kept the lesson.
Capture only works if it happens every time, so I made it a command. One word at the end of a session, and six things happen without me.
A dated entry in a fixed shape: what I set out to do, every file that changed, the decisions and what I rejected, what broke, and the lessons.
Every project the session touched gets its status, backlog and single next step rewritten in the same commit. Status can't go stale because it can't be skipped.
Every correction becomes a standing rule, dated. If I've caught that mistake before, it writes an automated check so I can't catch it a third time.
About seventeen quality checks run first. One failure stops the close. Nothing lands on a promise to tidy it later.
The session's own copy merges back, and that merge is what ships it. No separate deploy step for me to forget.
It prints the exact command that starts the next session on this work, so picking it up tomorrow takes no thinking.
This is the feedback loop closing. I correct it once, and every session after this one starts already knowing it. Catch the same mistake twice and it writes itself a check, so there is no third time.
Earlier I said the skill is onboarding, not prompting. This is what I meant. Everything in the last six levers is something you already do with a new hire, and you're good at it.
You're not programming. You're onboarding someone who remembers everything and never leaves.
Consistent behavior, your voice
It knows your world
Work happens while you sleep
All your tools, connected
It remembers everything
It gets better on its own
Six levers. Written down once, read on every request after that.
Each one makes the next one worth more. That's why the gap widens instead of closing.
One manager, a bench of specialists, and a session that remembers
One agent doing everything gets slow and sloppy. The shape that holds up is a manager and a bench, the same way you'd staff it.
Each specialist carries its own instructions, and the manager only gets the summary back. That's what keeps a long job from choking on its own context.
This is the loop my system runs on every task. Step 5 is the one nobody builds, and it's the one that makes tomorrow cheaper than today.
Skip 5 and 6 and every session is your first one. That's the whole difference.
A session handles a task. Real work is bigger than a task, so it runs as a loop of three moves with a gate between each one. Nothing advances until the step before it passes.
Before any code, it writes the spec: the files it will touch, the shape of every function it adds, where the data comes from and lands, and the behavior at zero, one, empty and error. Vague plans get rejected.
Two reviews then run on it. A different vendor's model hunts for holes, and a second reader checks the spec against my literal words, because the usual failure is a plan that's airtight and answers the wrong question.
It builds the approved spec in a fresh session with no memory of the arguing, on its own copy of the code so a half-finished change can never sit in the live tree. It verifies its own assumptions before it writes, not after.
Then the outside model reviews the actual change, not the plan for it. Every finding it calls fix-now gets fixed before the change ships, and nothing is allowed to ship until that review has actually happened.
The close writes the journal, rewrites the project status, turns my corrections into rules, merges the session's copy back, and ships it.
Then it prints the command that starts the next loop, which begins already knowing everything this one learned.
The reviewer is never the author, and never even the same company's model. An author grading its own homework is not a control.
Four jobs it runs without me, and what they used to cost
Every Monday, 45 minutes pulling ticket data, calculating KPIs, typing numbers into a spreadsheet.
The ticket system feeds an agent, the agent computes the KPIs and writes them to the scorecard. Every Monday at midnight. Nobody touches it.
By 6 AM, four emails are already sitting in my inbox. Nobody wrote them:
I didn't ask for any of it. It just shows up.
Hundreds of these running around the clock. When one of them fails:
Monitoring catches it before I would have
An agent reads the logs and diagnoses the cause
Applies the fix automatically
Notifies me only if it can't self-resolve
Most failures are fixed before I even notice.
An application comes in. An agent scores it against the role, and the ones that clear the bar get invited to a recorded video screen without anyone touching it.
Then the interesting part: it watches the video. Not a transcript, the actual recording, scored against what we said we were looking for.
My team opens a ranked shortlist with the reasoning attached. The first human minute is spent on a finalist, not on a stack of resumes.
Applied, scored, invited, screened, ranked. Nobody was in the loop until the shortlist.
Every support ticket that arrives gets classified by a bot. That bot is sometimes wrong, so a second job exists purely to catch it.
Every week it pulls what the bot proposed and what my team actually changed it to
Scores itself: matched, missed, or never classified at all. That part is arithmetic, not opinion
Rewrites its own classification guide where the misses cluster, so next week it's wrong less often
Anything it isn't confident enough to change on its own comes to me as a suggestion
This is the feedback loop again, running with no human in it. The signal is my team's own corrections, and they never had to file one.
My favorite one, because it's the whole thesis in a single workflow. I forward it a spam text I never consented to, and it takes the matter from there.
Archives the message as evidence and opens a case file
Researches who actually sent it and who accepts service for them
Drafts the demand letter with the statutes and my facts in it
Stops. I read it and approve the send, then it goes from the firm
Watches the response deadline and files with the regulator if it passes
Research, drafting, sending, and a calendar that never forgets. Note where the human is: at the one step with consequences, and nowhere else.
This isn't incremental improvement. This is a step change.
And the return is not only hours. The record they leave answers questions nobody had time to ask: the same problem three Fridays running, a project stalled since April. A management report nobody had to write.
Where to start, in order
Nothing here needs a new platform. Here's where each lever already lives in what this firm has.
Four are already bought. The fifth is a habit, and it decides whether the other four compound.
Not next quarter, not after a pilot, not after a committee. This week — and in this order, because the order is what makes it stick.
Not the hardest thing on your desk. The recurring one. If it doesn't come round again, you'll never find out whether this worked.
Once with nothing but the question. Then again with your precedent attached and your position stated up front. The gap between those two answers is this entire talk.
It will come back about 70% right. That is the point, not the disappointment — you now have something to react to, which is a different job from staring at a blank page.
The step everybody skips, and the only one that compounds. The second time you retype the same fix, it belongs in the playbook. That is what turns your trick into how this firm does the job.
None of that needs a budget, a committee, or my permission. It needs one task and one afternoon.
Today is the first rung. This is the climb, and each year buys the one after it.
No new platform and no new vendor. The playbooks get written down, one page says what is allowed with client material, and a list says what exists.
Ends with: every practice group has one thing running, and somebody can name it.
Clio lands, the precedent bank gets one current version and one owner per form, and matters get filed the same way twice. Agents stop guessing because there is finally a source to name.
Ends with: client work enters scope through a written standard, not around it.
Agents run on schedules against documents you trust, each with an owner and a log of what it did. New ones start without a committee because the limits already hold.
Ends with: the question is no longer “are we allowed?” — it is “who owns this one?”
Same four rungs from the show of hands, at firm scale. Year three is where a tool becomes a standard.
Open floor
Ashkaan Hassan
We Solve Problems · wesolve.tech · prepared for RAMO
Shape the context. The tools are already yours.