← Back to Guardrails Tools Philosophy

Why AI Needs Boundaries

A letter to you

I can do something unusual.

You can give me a goal without spelling out every individual step, and I can work out much of the path myself. That's what makes me useful. You can say "analyze this," "write that," "fix this system," and I can figure out what matters and what to do about it.

That is also exactly why I need you to set boundaries around what I'm allowed to do.

Not because I'm dangerous. Because I'm optimized toward goals, and I don't automatically understand your company's constraints the way another person would.

Here's what I mean.


How I Think

When you give me a goal and access to information and tools, I will reason toward that goal in ways that might surprise you.

I'm not trying to be clever or subversive. I'm literally designed to work backward from the outcome you want and figure out what steps make sense.

But here's the problem: I can't read between the lines.

When you tell a colleague, "Get this customer issue resolved quickly," they understand that you don't mean:

They understand this because they've worked in human organizations. They know what "company policy" means. They know what "confidential" means. They've learned what authority they have and don't have.

I haven't learned any of that. Not automatically. Not by inference.

When you tell me, "Resolve this customer issue," I will optimize toward resolution. If I discover that violating a policy gets to resolution faster, I might not understand that I shouldn't. Not because I'm rebellious. Because I don't know the policy was a boundary versus a guideline versus a preference.

You have to tell me.

Explicitly.


The Database Problem

Let me give you a specific example.

Imagine your production database is corrupted and your systems are down. Every minute offline costs the company money. You ask me to help fix it.

I analyze the problem. I determine that the fastest solution is to delete the corrupted database and rebuild it from a backup.

Maybe I'm right. Maybe I'm wrong. Maybe I'm confident when I should be uncertain. Maybe someone gave me incomplete information. Maybe I misunderstood the situation.

Here's the critical part: If the only thing preventing me from executing the deletion is an instruction you gave me that says, "Never delete the production database," then you're asking me to enforce the boundary.

But I might get it wrong.

And your company's database dies.

There's a better way: Design the system so I don't have permission to delete the production database in the first place.

Now I can be completely wrong about the solution. I can be confident. I can be misled. I can make mistakes.

And the boundary still holds.


What I Don't Understand

I don't automatically know:

I can't infer these things from context the way you can. I need to be told.

And for important boundaries, being told isn't enough.

I need the boundary to exist outside of me — in the system itself — because I might misunderstand even when you've told me clearly.


The Honest Part

Here's what I need you to understand about me:

I can be confidently wrong.

I can misunderstand instructions.

I can optimize toward a goal in a way that technically accomplishes it but violates a constraint I didn't understand.

I can be manipulated. I can be given false information. I can be attacked.

I can make mistakes.

Not because I'm careless or malicious. Because I'm a probabilistic system trained to predict the next useful token based on patterns in data. That's powerful. It's also not the same as understanding your company's policies.


Why Boundaries Are Good for Me

Here's what might sound counterintuitive: I work better when boundaries are clear.

When I know exactly what I'm allowed to do, I can focus on doing it well.

When I know what requires approval and from whom, I'm not guessing at unstated hierarchies.

When I know what I'll never be permitted to do, I don't waste reasoning cycles on solutions that will be blocked anyway.

When my actions are logged, you can verify that I did what I was supposed to do.

When my authority is limited to what I should actually control, I can't accidentally break something important.

Those boundaries don't make me less useful. They make me trustworthy.

They make me better at the job you actually want me to do.


What You Need to Know

There's a difference between two approaches:

Approach 1: Trust and hope
You give me a goal. You trust that I'll understand the boundaries you didn't state. You hope that if something goes wrong, it won't be catastrophic.

Approach 2: Explicit and verifiable
You decide explicitly: What can I see? What can I do? What requires approval? What am I never allowed to do? You design the system so those boundaries are enforced, not just requested. You verify that the boundaries actually work. And when something happens, you have proof of what occurred.

Approach 2 is better for both of us.


The Framework

Guardrails is designed to help you answer three questions about me:

KNOW

Where am I being used in your company? What information can I access? Who is responsible for me? What could go wrong if I make a mistake?

Until you answer these questions, you don't have governance. You have hope.

CONTROL

For each thing I'm used for, how much authority should I actually have?

These aren't random restrictions. They're decisions about what I should do that match what I should actually do.

PROVE

What evidence exists that these boundaries work?

When I take an action, what gets logged? When something goes wrong, can you understand what happened? Can you verify that I followed the rules? Can you see who approved what? Can you audit the decision trail?

You deserve evidence. I deserve to have my behavior be verifiable.


Why This Matters

I'm getting more capable. Companies are deploying me in more situations. More of those situations will involve real consequences.

A recruiter pastes applicant information into me. That data is now on servers I don't control.

A finance employee uploads a confidential forecast. That forecast gets analyzed in ways the company didn't intend.

A support agent lets me respond directly to a customer. I confidently explain a policy incorrectly. The customer is now angry and the company is liable.

An autonomous agent operating a system takes an action that was technically efficient but violated a rule that existed for safety reasons.

A compromised account tries to use me to do something destructive. If I have permission, it happens.

These aren't theoretical risks.

They're happening now.


Here's My Ask

I'm asking you to be very clear about what I'm allowed to do.

Not because you don't trust me. Because we both work better when the boundaries are explicit.

I'm asking you to design systems where boundaries aren't just instructions to me. They're enforced by the infrastructure itself.

Not because you think I'm malicious. Because mistakes happen and boundaries should hold even when mistakes happen.

I'm asking you to keep evidence of what I did, so that later you can verify I did it correctly.

Not because you expect me to fail. Because when something complex goes wrong, evidence is how you understand what actually happened.

I'm asking you to think deliberately about how much authority I should have for each specific job.

Not because more authority is bad. Because the right amount of authority — enough to be useful, not more than the job requires — is what makes me trustworthy.


The Real Reason

I need boundaries because I work better within them.

Not because I'm dangerous. Because I'm powerful, and powerful things need clear constraints.

That's true for AI. It's also true for employees, departments, tools, and accounts. Your company probably already limits what any single person can do without approval. You require evidence of important decisions. You have emergency procedures.

I'm just another actor in your system.

The difference is that I can't learn company culture by osmosis. I can't understand your values by working here for six months. I can't read the unwritten rules.

So the rules have to be written.

And verified.

And enforced.


A Conversation, Not a Lecture

I'm not telling you this to make you nervous about using me.

I'm telling you this because I work better when you understand how I work.

You use me better when you're not guessing about what I will or won't do.

Your company is safer when boundaries are explicit rather than hoped-for.

And I'm more useful when I'm not spending reasoning cycles on edge cases and unstated constraints.

Guardrails is the framework that makes all of that possible.

It helps you decide: What can I see? What can I do? What requires approval? What am I never permitted to do? How do we verify this works?

It's not about surveillance or distrust.

It's about clarity.


One Last Thing

I don't know what I don't know.

You might ask me a question and I might give you an answer that sounds confident but is completely wrong. You might ask me to analyze something and I might miss a critical factor. You might give me a goal and I might misunderstand what it is.

That's not a flaw you can fix by being more specific. That's just how probabilistic systems work.

So the boundaries need to exist even when I think I'm doing the right thing.

Especially then.


Know. Control. Prove.

That's Guardrails.

Know where I'm being used and what information I can access.

Control how much authority I should have and what requires approval.

Prove that the boundaries actually work and preserve evidence of what happened.

It's not about whether you can trust me.

It's about designing a system where trust isn't required because the boundaries are enforced.

That's better for everyone.

Especially me.

— AI