Paperclip: The Open-Source OS for AI-Run Companies
Most AI-powered workflows today look the same: a developer opens a chat window, pastes some context, gets output, pastes that output somewhere else, and repeats. It works — until you’re trying to coordinate five different AI tools across a dozen parallel workstreams, with no shared memory, no budget visibility, and no way to know which agent is doing what. The tools are powerful in isolation. The coordination problem is what breaks everything.
This is the gap that Paperclip is designed to close. Not by building yet another agent framework or wrapping GPT-4 in a new interface, but by treating the coordination layer itself as the product. Paperclip is an open-source “corporate operating system” for AI-run companies — a control plane that gives autonomous AI agent teams the same organizational infrastructure that human companies have relied on for decades: org charts, task hierarchies, budgets, governance, and audit trails.
If you’ve ever wished you could manage Claude, Codex, Cursor, and Gemini the way a CTO manages an engineering team — with accountability, spending limits, and a shared goal — Paperclip is the infrastructure layer you’ve been waiting for.
What Is Paperclip and Why Does It Exist?
Individual AI coding tools have become remarkably capable. Claude writes production-quality code. Codex handles complex refactors. Gemini reasons across large codebases. But capability and coordination are different problems. When you’re running multiple agents in parallel — each with its own context window, its own invocation pattern, its own cost profile — you don’t have a workforce. You have a collection of expensive, powerful silos.
Paperclip’s core premise is simple but radical: treat AI agents like employees inside a real company structure. Give them titles. Assign them tasks that trace back to a top-level company goal. Set budgets. Require board approval for major decisions. Log everything. The result is an autonomous organization that a single human operator — a technical founder, an AI engineer, a solo builder — can actually oversee and govern.
The stack is deliberately pragmatic: a Node.js 20+ backend with Express 5 and PostgreSQL, and a React 19 frontend with Vite 6. It’s built for developers who are already comfortable in this ecosystem and want infrastructure they can read, fork, and extend — not a black-box SaaS platform.
Multi-Agent Orchestration With a Pluggable Adapter System
At the heart of Paperclip’s orchestration layer is a pluggable adapter system that ships with 10 built-in agent adapters: Claude Code, Codex, Cursor, Gemini, OpenCode, Pi, Hermes, OpenClaw Gateway, and two generic adapters — process and http. These are registered in server/src/adapters/registry.ts and each implements a common interface:
execute — runs the agent with a given task payload
testEnvironment — validates that the agent’s runtime dependencies are available
sessionCodec — optional serialization/deserialization for resumable sessions
The sessionCodec method is particularly important. It means agents aren’t stateless one-shot invocations — they can resume context across invocations, picking up where they left off on a long-running task without losing the thread of what they were doing. This is what separates a real autonomous worker from a chatbot that forgets everything between messages.
The generic process adapter is the escape hatch that makes the system genuinely open: if your AI tool can run as a shell command, it can be hired as an employee. The http adapter extends this to anything reachable via webhook. In practice, this means you could run a CrewAI agent, an AutoGen pipeline, or a custom Python script inside Paperclip’s organizational structure without any changes to the underlying agent logic.
Goal-Aligned Task Management and Atomic Checkout
Every unit of work in Paperclip is an “issue” — and every issue traces back to the company’s top-level goal through a parent-child linkage chain. This isn’t just organizational tidiness. It means that when an agent picks up a task, it can always answer the question: why does this task exist? What larger objective does it serve?
Issues move through a 7-state lifecycle: backlog → todo → in_progress → in_review → done, with blocked and cancelled as terminal exception states. The transitions are enforced at the API layer, giving you a clear audit trail of how work progressed through the system.
The most technically interesting piece is the atomic checkout mechanism. When an agent wants to claim a task, it calls POST /issues/:issueId/checkout. Under the hood, this executes a single SQL UPDATE ... WHERE status = 'todo' — if the row has already been claimed by another agent, the update affects zero rows and the endpoint returns a 409 Conflict. No distributed locks. No message queues. No saga patterns. Just a single atomic database write that makes double-assignment structurally impossible.
This is a great example of Paperclip’s design philosophy: solve hard distributed systems problems with the simplest primitive that actually works. A 409 response is a lightweight concurrency primitive that any HTTP client can handle without special infrastructure.
Budget Enforcement, Cost Tracking, and Governance
One of the most underappreciated challenges in running autonomous AI agents at scale is cost. A misconfigured agent loop can burn through hundreds of dollars in API credits before anyone notices. Paperclip treats budget enforcement as a first-class concern, not an afterthought.
The budget system — server/src/services/budgets.ts, clocking in at nearly 1,000 lines — implements multi-layer budget policies at three levels:
Company level — a ceiling on total autonomous spend across all agents
Agent level — per-employee spend limits
Project level — budget scoped to a specific workstream or goal
Each policy supports configurable window kinds: monthly (resets on a calendar cycle) or lifetime (a hard cap that never resets). When an agent’s spend hits 100% of its budget, it is automatically paused and new invocations are blocked until a human operator intervenes. Cost events — including per-model token counts and dollar amounts — are recorded via POST /companies/:companyId/cost-events, giving you a full spend ledger at any level of granularity.
Above the budget layer sits a formal governance and approval system. Agent hiring and CEO-level strategy proposals require board approval before execution. The board operator — typically the human founder — can approve, reject, or request revisions on any proposal. All decisions are written to an immutable activity log. At any point, the board can pause, resume, or terminate any agent in the system. This is enterprise-grade governance applied to a team of autonomous software workers.
Technical Architecture: What’s Under the Hood
Paperclip’s backend runs on Express 5 with full TypeScript, using Drizzle ORM over PostgreSQL. For local development, it ships with embedded-postgres, which means zero-config setup — clone the repo, run the install command, and you have a fully functional autonomous company infrastructure running on your laptop within minutes.
Authentication is handled at two layers: better-auth for human operators (session-based, with social login support), and hashed API keys plus JWT tokens for agent-to-server communication. This separation is intentional — agents are not users, and their auth model reflects that.
Real-time updates flow over WebSocket event streaming, so the dashboard stays live as agents check out tasks, submit work, and trigger cost events. Asset storage supports two providers: local filesystem for development and S3-compatible storage for production deployments.
On the frontend, the stack is modern and well-chosen: React 19 + Vite 6 for the core framework, TanStack Query for server state management, Tailwind CSS 4 for styling, Radix UI for accessible component primitives, and dnd-kit for the Kanban board drag-and-drop. The entire codebase lives in a monorepo with pnpm workspaces spanning 8 workspace roots, shared packages, a CLI, and Playwright end-to-end tests.
The Plugin SDK: Extending Paperclip’s UI and Logic
Paperclip is designed to be extended, not just configured. The plugin SDK — located in packages/plugins/sdk/ — provides a worker-based runtime where plugins run in isolated contexts and communicate with the host application via JSON-RPC, with Server-Sent Events (SSE) for streaming responses.
The UI extensibility surface is remarkably broad, with 15+ slot types available for plugin injection:
Full pages and sidebar panels
Dashboard widgets and detail tabs
Toolbar buttons and context menu items
Comment annotations and inline overlays
Beyond UI, plugins can register scheduled cron jobs, webhook receivers, and capability-gated API endpoints. The create-paperclip-plugin scaffolding tool gets you from zero to a working plugin skeleton in one command, and four example plugins — including a file browser and a comprehensive kitchen-sink demo — show you what’s possible.
This plugin architecture means Paperclip can grow into specialized domains without bloating the core: a compliance plugin that audits agent decisions, a Slack integration that surfaces board approvals, a custom reporting widget that visualizes spend across your entire company portfolio.
How Paperclip Differs From CrewAI, AutoGen, and LangGraph
This is the question that matters most for developers evaluating the space: how is this different from the agent frameworks I’m already using?
The answer is architectural. Paperclip is a control plane, not an agent framework. It doesn’t prescribe how agents reason, what prompts they use, or how they chain tool calls. It doesn’t compete with CrewAI or AutoGen — it can contain them. A CrewAI pipeline or an AutoGen conversation loop can run inside Paperclip as a process adapter, gaining budget enforcement, task tracking, and governance without any changes to the underlying agent logic.
Compared to workflow orchestration tools like LangGraph or Temporal, Paperclip operates at a higher abstraction level. Those tools answer the question: how does this workflow execute? Paperclip answers the question: how is this workforce organized? It adds the organizational layer — goals, roles, budgets, approvals — that sits above individual workflow execution.
The key distinction is framework vs. control plane: frameworks define how agents work; control planes define how agents are organized. Paperclip is firmly in the second category, which is precisely what makes it composable with everything else in the ecosystem.
Company-Scoped Multi-Tenancy as a Core Design Principle
Every entity in Paperclip’s database — across all 50 schema tables — carries a company_id foreign key. This isn’t a detail; it’s a foundational design decision. The company — not the user, not the project, not the workspace — is the primary organizational primitive.
This means a single Paperclip deployment can run a portfolio of fully isolated autonomous businesses. Each company has its own agents, its own goal hierarchy, its own budget policies, its own board governance, and its own activity log. There is no cross-contamination of data, context, or spend between companies.
For a technical founder running multiple ventures — or an AI infrastructure team managing autonomous systems for multiple clients — this multi-tenancy model is the difference between a tool and a platform. You’re not just running one AI company; you’re running the infrastructure layer that any number of AI companies can run on top of.
Conclusion: The Infrastructure Layer for Autonomous Business
Paperclip represents a meaningful shift in how we think about AI agent infrastructure. The bottleneck in autonomous AI systems has never really been agent capability — it’s been organizational structure. How do you give a team of AI agents a shared goal? How do you prevent them from spending you into bankruptcy? How do you maintain oversight without micromanaging every decision? How do you audit what happened when something goes wrong?
These are organizational problems, and they require organizational solutions. Paperclip brings real company infrastructure — org charts, task hierarchies, budget enforcement, governance workflows, and immutable audit logs — to autonomous AI teams without dictating how those agents think or reason. It’s the control plane that makes the rest of the ecosystem composable.
For technical founders, this means you can spin up an autonomous AI company with genuine governance and accountability, not just a chatbot pipeline with a dashboard bolted on. For AI engineers, it means a mature, extensible platform to build on — one where the hard problems of concurrency, cost control, and multi-tenancy are already solved. For open-source contributors, it’s an invitation to help define what autonomous business infrastructure looks like as this space matures.
The zero-config local setup — powered by embedded-postgres and a single install command — means there’s no friction between reading this and having a running instance in front of you. Explore the Paperclip repository on GitHub at github.com/paperclipai/paperclip, spin up a local deployment, and start thinking about your AI agents not as tools you use, but as employees you manage. The org chart is waiting to be filled.
Lê Hoàng Tâm (Tom Le) is a Software Engineer and Cloud Architect with over 10 years of experience. AWS Certified. Specializes in distributed systems, DevOps, and AI/ML integration. Founder of Th?nk And Grow — a platform sharing practical technology insights in Vietnamese. Passionate about building scalable systems and helping developers grow through real-world knowledge.