Access Control · Internal platform

Adding a permission shouldn't mean altering the user table

Permission models in internal tools usually begin as a checkbox on the user record, and that holds up until there are forty screens. By then the shape is load-bearing: adding a permission means altering the table every account lives in, and the question of who can open a given screen has no answer except reading the code of the screen. Nobody rewrites that while the tool is in daily use — which is precisely why it needs rewriting.

55 operational screens
2 ranked privilege tiers
17 allowlisted unauthenticated calls
Confidential client & operation

Built for the internal operations platform of a contract-logistics site, where floor operators, analysts, leads and site management all work from the same sign-in.

The problem

Permissions lived as bit columns on the user row. The application discovered what roles existed by querying the database's own schema catalogue for every boolean column on that table — so a role was a column, granting one was a bit flip, and creating one was a schema change to the table every account depends on. Nothing anywhere recorded which screens a role actually opened. That mapping existed only inside whichever page remembered to check the flag.

Credentials were worse. Passwords were stored under a reversible letter-shift cipher, which is an encoding and not a hash — anyone who could read the table could read every password in it. Both problems had the same root cause, which is that the model was never designed; it accreted.

  • A new permission required a schema change to the user table
  • The list of roles was derived at runtime by reading the database schema
  • No record of which screens a role granted — only which flag a page happened to test
  • Stored passwords were recoverable by anyone who could read the table

What we built

Three tables and one rule. A role is a name plus a set of screens; a user holds roles; a screen is a row in a catalogue. Granting access became inserting a row rather than altering a schema, and deleting a role cascades to both its screen grants and its user assignments in the same statement — there is no such thing as an orphaned permission left behind.

The permission tables create themselves on first use and add missing columns additively. That is how the second privilege tier was introduced to a system already in daily use, without a migration window or a coordinated deploy.

Page permissions are resolved once, at sign-in, into the session. Every check after that is a set-membership test with no database round trip, because an authorization check on the request path has to be effectively free — the moment it isn't, people start writing code that skips it. The cost of that choice is honest and bounded: a permission change takes effect at the user's next sign-in, and sessions expire after four hours idle.

  • Roles are declarative — saving a role replaces its entire screen list rather than patching it
  • Registering a new screen is an idempotent insert into the catalogue, not a schema step
  • The catalogue is cached for five minutes, so navigation doesn't re-query it on every page load
  • An expired session returns JSON to an API caller and a redirect to a browser, so a background fetch never has to parse a login page
  • The stored password hash is stripped from the user record before anything is written to the session
  • Kiosk screens that run unattended reach the server through an explicit allowlist of seventeen module-and- function pairs; that module's administrative functions stay behind a separate leadership check

Proving it's a manager

Some actions need a second person, not a more privileged session. Removing an expected pallet from a load audit, overriding a scan discrepancy, forcing a failed line to pass — the operator does the work, a manager authorizes the exception, and the operator carries on as themselves. Elevating the operator's session would be the easy implementation and the wrong one, because the session is what the audit trail is written against.

So the credential check is a separate endpoint that validates a username and password through the same path as sign-in and then deliberately does nothing else. No session is created, none is modified, no privilege is handed to the browser that asked. It answers two things and nothing more — whether the credentials are valid, and whether that account holds the tier the action requires.

The two tiers are ranked, and the ranking runs one way only. A manager passes a leadership check, because a manager directs the leads. Leadership credentials do not pass a manager check. Getting that asymmetry wrong in either direction produces the kind of bug nobody finds until an override has already been authorized by someone who shouldn't have been able to.

Every authorization carries a written reason. The prompt refuses to submit without one, and the reason is stored on the record with the authorizing username and the timestamp. One fix in that area is worth naming: a later recalculation pass used to rewrite overridden rows back to a normal status and drop the badge marking them, which erased the evidence that a person had intervened at all. Overridden rows are now excluded from that pass explicitly.

Failing closed

The platform can also be suspended wholesale, out of band. A flag file at the application root flips a request hook that answers every route with a suspension page — HTML for browsers, JSON for API paths, so a client-side fetch gets something it can actually read.

The flag is a file rather than a database row, because the check runs ahead of every single request and must never cost a round trip; each process re-reads it at most once every five seconds. It is written by atomic replace, so a half-written file can never be observed mid-write. If the file exists but cannot be parsed, the service stays suspended — a corrupt write must not silently restore access. The hook is registered ahead of the session hooks so nothing else runs, and the metrics endpoint is left open on purpose, so monitoring reports an outage rather than a healthy silence.

The instruction channel is a polled mailbox, and it is treated as one rather than trusted as one. Each message is applied exactly once, keyed on its message identifier and recorded in an audit table. Pending messages are applied oldest-first, so when two contradictory instructions are waiting the most recent one decides the final state. An optional sender allowlist rejects anything else and records the rejection. The poll query is scoped to the inbox so the system's own confirmation replies cannot re-trigger it, and only one process in the fleet runs the poll — the others learn the state from the file.

The result

Permissions became data. A role is a row, a grant is a row, and what a person can open is a query rather than a boolean column plus a page that remembers to check it. The password store was migrated in place to a proper hash, without asking anyone to reset anything and without a flag day — the migration skips rows already converted, so it can be run again safely.

The parts we would defend hardest are the unglamorous ones. That a privileged action verifies a second person without elevating the first. That every override carries a name and a written reason that later processing cannot quietly overwrite. And that the mechanism capable of turning the whole application off fails toward off.

Built with

  • Python
  • Flask
  • SQL Server
  • bcrypt
  • Gmail API
  • Server-rendered templates

On the details we left out

Client, location, system names and volumes are withheld deliberately. We're happy to go deeper on architecture and approach on a call.

Services this draws on

Legacy Modernization & Migrations

Bring decades-old systems into the modern era without a risky rip-and-replace.

See the service →

Custom Software & Web Development

Websites, internal tools, APIs, and automation for teams without a dev department.

See the service →

More projects

Product · Quality & compliance

Audits that can't be quietly passed

Customer-facing quality audits ran on paper and a survey tool. We built a configurable audit platform with scheduling, escalation, approvals — and a gate that blocks sign-off when the scan doesn't match.

Read the story →
Safety · Equipment control

Equipment that locks itself out when it fails inspection

Powered industrial trucks were signed out on a clipboard, and the fall-protection harness that some of them require was a separate honour system. We built a badge kiosk where the truck you pick decides the inspection you must pass — and which refuses to release anything that fails.

Read the story →
Controls · Reverse logistics

Financial controls on material leaving the building

End-of-life material left the site on paperwork typed by hand. We rebuilt the chain end to end and added the control nobody had: detection of inventory added after the paperwork was already filed.

Read the story →

Got a system like this?

A free assessment gives you a clear picture and a prioritized plan — no commitment.

Book a free assessment