https://github.com/SenteLabsAI/OpenExecutive/blob/main/packa...
βWalking away costs you the deal. Signing costs you the company. The math is obvious in retrospect; harder in the moment when there's quota pressureβ
such valuable, should I say load-bearing insights
Other gem from my new CFO:
βThe financial statements are the result. The variance analysis is where the operating insight lives.β
And my new C(3-)PO just told me
βIf we removed all sales involvement, what % of revenue would we still get? #5 is the honest PLG test. If the answer is <30%, PLG is supplementary at best. If the answer is >60%, PLG is the actual motion and the rest is overlay.β
Imagine how quickly your meetings could get resolved if you had teams submit prompts / contexts, had a clear, documented set of overarching objectives, etc.?
That said, if you look at what itβs actually doing (read the question carefully, then read what it actually said) I am not convinced this is better than what ChatGPT could do out of the box already.
Sometimes the claim is "set vision" - which AI can extract from the collective.
Someimes the claim is to prioritize the work (with bias), AI can extract that from the collective too, right?
Making resources available to the team? AI.
Coordinating between teams? AI.
Can the AI do this in a more un-biased way? Can it refine collective intelligence and desire into objective goals?
More than a software factory id love to see a manager-factory.
One thing, AI cannot do (yet) is bring the network. Like when you hire folk from Booth or Sloan so you can sell to places that hired Booth and Sloan grads.
The thought experiment, to be clear, is building an artificially intelligent CEO or company founder with the intent of generating profit. If this AI can turn a profit, it has passed this new test
Of course, by the time this new AI starts turning a profit, it will be maximized to the extent that it loses its competitive advantage, and then weβre back to square 1. That is, when everyone uses the same AI, nobody has a competitive advantage over anyone else
My current theory is that because of this, it is not possible to build an AI that generates a profit perpetually. Only temporarily. Then human intervention would be required to adjust things to make the AI competitive again. But at that point, itβs no longer an AI driving the profit, itβs its human tinkerer
Interested folks can start by putting in shareholder resolutions with the proposal and the prompts that would replace CEO. SKILLs.md?
Please replace CTOs while you are at it and most architects can also be replaced easily!
A very personal view but many developers/engineers fail to see the overall picture. And yes, I was there and I went through all of the mistakes.
The things you see on the surface of those incompetent CxOs are just the tiny piece of the bigger picture. You cannot replace blames, decision in the face of uncertainty. Almost all decisions that businesses take have nothing to do with the proficiencies of technologies or engineering finesse, or even the best options, but mostly on βincentivesβ and who can take/own the βblame.β
βAIβ replacing developers is evident because that is a rather tiny perspective and kinda easy to do the equation - developer can do that (good, bad, best) and can the AI do it - maybe, perhaps. So, replace.
A C-Suite took a decision to tick the final approval of a billion dollar deal with a service providers (amongst 10s of them). The idea might have sounded idiotic because there were 3 others better at the top than the chosen one. Even if all of the task and options were weighted with AI, but that last human decision is the Blame that the C-Suite took on, by betting his incentives on it (money, fame, pride, ego.)
The way any large company works, and this includes all the big tech companies, is that they end up in a state of constant churn. Every 6 months or so there's a reorganization. There'san old joke that you get an email that your manager's manager's manager's manager now reports to a new manager and you've never heard of any of these people.
The point of that churn is to make sure nobody is ever held responsible for something not working. It's never given enough time. That's by design. Middle (and even senior) management is about empire-building.
Another quote that springs to mind is that the bureaucracy is expanding to meet the needs of the expanding bureaucracy. That's how I feel about management bloat.
Anyone who has paid attention to Intel's downfall in the last decade will have heard stories about how the company descended into fiefdoms and it became impossible to get anything done. That's what happens.
Most people in these organizations aren't particularly talented at anything other than forming social relationships. the leadership of these companies always end up having an adversarial relationship with their workers and try to suppress their wages to increase profits so they can get their stock options and bonuses and the CEOs all sit on each other's boards of directors setting each other's compensation packages. It's a lootocracy.
The point of all this is that the leadership will never decide that they are the problem. Take Meta's Metaverse. It was a $70 billion flop. A lot of people got laid off from that. Were any of those worker bees responsible? Nope. They didn't choose the strategy or probably even what they worked on. It takes a lot for heads to roll within senior leadership and with the vote-stacked share structures of many major tech companies, ousting a CEO is essentially impossible.
Be friends with the board. Negotiate high pay and benefits (for themselves), and a golden parachute in case things go south (which they do most often). Take credit for things that were going to happen anyways. Blame xyz for things that didn't work out.
These are all impossible tasks for AI. AI can't play golf.
The answer is to be responsible to the business. The buck stops at the CEO. Therefore, their responsibilities are technically everything in the entire company. And I think many people who have only worked for other people underestimate how valuable that responsibility is.
> The job of an executive is: to define and enforce culture and values for their whole organization, and to ratify good decisions.
(And that's the whole list! Nowhere do you find set tasks or make decisions)
To me, that means that they're still pretty essential - on both fronts, AI needs values and AI needs a system that can tell whether decisions are good or not.
That is what the CEO is for β to be on the receiving end of beatings and flogging if the business they are at the masthead of effs something up.
To be an effective CEO, you have to make judgement calls that the collective canβt always have wisdom and context on, is confidential, or requires fast decision making during which there are interactive discussions or negotiations!
This argument always confuses me because it seems like lower level employees are the ones that get quickly dismissed when things get bad while a lot of C level executives seem to have some sort of compensation even if they fuck up.
Sun Tzu said to "know your enemy" and those who don't know the basics of how businesses and finance work have no hope of defeating them.
People always dunk on me for saying this outrageous statement, but if you put me in charge of a random Fortune 500 company, I give myself a 50% probability of the company performing better than if its real CEO was in charge, and a 50% probability of it performing worse. I really don't think these guys bring as much to the table as they think they do and as much as the general public seems to think they do. Business success does not come from paranormal insight and abilities.
I think you could randomly nominate CEOs from the set of the general population with, say, a minimum of a college degree, and the randomly nominated CEO's performance would not statistically differ from "qualified" CEOs.
What is failling to see the overall picture? I hate this overall picture mysticism. I bet that if people where given the same role and information as the CEO they would be able to see the overall picture.
Most importantly, what is its golf handicap?
1. https://en.wikipedia.org/wiki/Tendency_of_the_rate_of_profit...
Well, so do most CEOs.
Unbounded liability means nobody starts a business, or they do it under the table without responsibility or taxes.
Turns out we want successful private businesses much more than we want to hold a few nasty CEOs accountable.
An AI system that acts as your company's virtual executive team β a senior advisor with Harvard MBA-level knowledge, customized for your specific business.
A walkthrough of Open Executive in action β watch on YouTube.
Developed by sentelabs.ai Open Executive provides a single coherent executive voice backed by eight specialist AI agents:
All responses come from one consistent executive voice. The internal agent architecture is never exposed to the user. Beyond Q&A, the system maintains episodic memory of past decisions and initiatives across sessions, and a built-in scheduler can proactively surface follow-ups and time-sensitive actions.
User message
β
Executive Orchestrator (claude-sonnet-4-6)
β tool use β parallel specialist calls
CSO / CFO / CHRO / GC / COO / CMO / CPO / Board
β each specialist retrieves relevant context from ChromaDB
Built-in MBA knowledge + Your company documents
β
Synthesized executive response
Knowledge β Two retrieval layers per specialist call: (1) built-in MBA-level Markdown (knowledge/builtin/, git-tracked) seeded into ChromaDB at startup, and (2) your uploaded company documents chunked and stored in a separate company_docs collection. RAG context is injected into the user turn, never the cached system prompt.
Episodic memory β After every response, a background claude-haiku-4-5 pass extracts key decisions, initiatives, and advice into SQLite. The next session opens with a <past_decisions> block so the Executive remembers what it recommended last month.
Scheduler β A built-in job runner claims due actions via UPDATE β¦ RETURNING to prevent double-firing. The API must run as a single instance; do not horizontally scale it without gating the scheduler first.
Prompt caching β The system prompt is structured so the Executive persona, company profile, and knowledge index are cached separately (up to 85% cache hit rate after the first few turns). No dynamic content ever goes in a cached block.
See docs/architecture.md for the full design.
| Layer | Choice |
|---|---|
| LLM backbone | Anthropic Claude API |
| Default model | claude-sonnet-4-6 (Executive + most specialists) |
| Deep reasoning | claude-opus-4-7 (CSO, CFO, GC, Board β with extended thinking) |
| Backend | Python 3.11 + FastAPI |
| Package manager | uv |
| Vector store | ChromaDB (local, embedded) |
| Episodic memory | SQLite |
| Web UI | Next.js 15 (App Router) + Tailwind |
| License | Apache 2.0 |
openexecutive/
βββ packages/
β βββ core/
β β βββ openexecutive/
β β βββ orchestrator/ # Executive persona + routing loop
β β βββ agents/ # 8 specialist agents
β β βββ knowledge/ # ChromaDB store + RAG pipeline
β β βββ memory/ # Company profile + episodic memory
β β βββ onboarding/ # Wizard state machine + profile builder
β β βββ prompts/ # Persona + domain prompts + cache manager
β β βββ api/ # FastAPI app + routes
β β βββ integrations/ # Slack, Email, Telegram, Google Chat, Discord
β β βββ scheduler/ # Background job runner (single-instance)
β β βββ alerts/ # Proactive alert system
β β βββ audit/ # Audit logging
β β βββ architecture/ # Internal architecture utilities
β β βββ workflows/ # Multi-step workflow definitions
β β βββ cli.py # Click CLI
β βββ ui/ # Next.js 15 web UI
βββ evals/ # Eval scenarios + LLM-as-judge runner
βββ fixtures/ # Demo company fixtures (profiles, docs, rosters)
βββ scripts/ # Operator scripts (Fly secrets, Google auth)
βββ docker/ # Dockerfile(s) + docker-compose.yml
βββ fly.api.toml / fly.ui.toml # Fly.io configs β dev API + UI apps
βββ fly.api.qa.toml / fly.ui.qa.toml # Fly.io configs β QA API + UI apps
βββ fly.honcho.toml # Fly.io config β Honcho memory app (optional)
βββ docs/ # Architecture + deployment docs
# Clone the repo
git clone https://github.com/SenteLabsAI/OpenExecutive.git
cd OpenExecutive
# Set your Anthropic API key
cp .env.example .env
# Edit .env and add ANTHROPIC_API_KEY=sk-ant-...
# Start everything
make dev
Open http://localhost:3000 to start chatting with your executive. The API runs on port 8000 and the UI on 3000.
First run: requires Python 3.11+ and Node 22+. The initial
uv syncpulls heavy ML dependencies (ChromaDB + sentence-transformers/PyTorch), and the first boot downloads a small embedding model (~90 MB) to build the local vector index β so the firstmake devtakes a few minutes before the app is ready. Subsequent starts are fast.
For contributors not using make:
cd packages/core
uv sync
source .venv/bin/activate
uvicorn openexecutive.api.main:app --reload --port 8000
# In a second terminal
cd packages/ui && npm install && npm run dev
bot + applications.commands scopes.env: DISCORD_BOT_TOKEN, DISCORD_APP_ID, DISCORD_GUILD_IDSDISCORD_BOT_TOKEN is set:make dev
The bot is embedded in the API process (alongside the email poller, scheduler, and resumer) so it shares the same SQLite database and ChromaDB vector store under /data in production. Skip the token to disable.
For iterating on bot-only code without restarting the API, make discord runs the bot as a standalone process against the same local DB.
Users can DM the bot, @mention it in a channel (replies in a thread), or use /ask and /today slash commands. Slash commands sync to DISCORD_GUILD_IDS instantly on startup; leave blank for global registration (up to 1-hour propagation delay).
Just set the secrets on the existing API app β no new Fly app required:
flyctl secrets set -a openexec-api-dev \
DISCORD_BOT_TOKEN=... \
DISCORD_APP_ID=... \
DISCORD_GUILD_IDS=...
Discord user access is managed via the /people UI β add a Person row with discord_user_id set.
The machine restarts and the bot starts on the next lifespan boot. To disable in prod: flyctl secrets unset -a openexec-api-dev DISCORD_BOT_TOKEN.
The first time you visit the app, you'll be guided through a wizard to set up your company profile:
After onboarding, the Executive will reference your specific company context in every response.
| Interface | How to Use |
|---|---|
| Web UI | http://localhost:3000 |
| Slack | Mention @OpenExecutive or DM the app |
| CC or email the configured address (IMAP/SMTP poller) | |
| Telegram | Message the configured bot |
| Google Chat | Mention the app in a space |
| Discord | DM the bot, @mention it in a channel, or use /ask / /today slash commands |
| CLI | openexecutive chat |
Upload your pitch deck, financial model, strategy docs, or any company documents via the web UI or API. The Executive will reference them when relevant.
# Via CLI
openexecutive upload deck.pdf model.xlsx strategy.md
# Via API
curl -X POST http://localhost:8000/documents \
-F "file=@deck.pdf" \
-F "domain=strategy"
Two environments, each a separate set of Fly apps, driven by branch:
| Environment | Trigger | Workflow | Apps |
|---|---|---|---|
| dev | push/merge to main (continuous) |
.github/workflows/deploy.yml |
openexec-api-dev, openexec-ui-dev |
| qa | push/merge to qa (deliberate promotion) |
.github/workflows/deploy-qa.yml |
openexec-api-qa, openexec-ui-qa |
Both workflows use dorny/paths-filter to deploy only the changed app (API, UI, or both). QA is a stable twin of dev β same image and runtime, only the app name differs (fly.api.qa.toml / fly.ui.qa.toml) β so it lags main and stays vetted. An optional Honcho memory app (fly.honcho.toml) deploys independently.
| App | Purpose | State |
|---|---|---|
openexec-api-{dev,qa} |
FastAPI + scheduler | Persistent volume executive_data at /data |
openexec-ui-{dev,qa} |
Next.js 15 | Stateless |
openexec-honcho-dev |
Honcho per-person memory (optional) | Postgres-backed |
β οΈ Single-instance only: The scheduler claims rows via
UPDATE β¦ RETURNING. Running two API machines would double-fire scheduled actions.max_machines_running = 1is set infly.api.toml/fly.api.qa.tomlβ do not override it.
Deploys authenticate with per-app Fly deploy tokens stored as repo (or org) Actions secrets. Generate each with flyctl tokens create deploy -a <app> -x 999999h:
| Secret | App | Used by |
|---|---|---|
FLY_API_TOKEN_API |
openexec-api-dev |
dev |
FLY_API_TOKEN_UI |
openexec-ui-dev |
dev |
FLY_API_TOKEN_HONCHO |
openexec-honcho-dev |
dev (honcho job) |
FLY_API_TOKEN_API_QA |
openexec-api-qa |
qa |
FLY_API_TOKEN_UI_QA |
openexec-ui-qa |
qa |
Per-app runtime secrets (ANTHROPIC_API_KEY, BACKEND_SHARED_SECRET, the AUTH_* set, integration tokens) are set directly on each Fly app β see scripts/fly-secrets.sh.example.
# 1. Create apps and volume
flyctl apps create openexec-api-dev
flyctl apps create openexec-ui-dev
flyctl volumes create executive_data --region iad --size 1 -a openexec-api-dev
# 2. Set the required secret
flyctl secrets set -a openexec-api-dev ANTHROPIC_API_KEY=sk-ant-...
# 3. Create deploy tokens and add as GitHub secrets FLY_API_TOKEN_API and FLY_API_TOKEN_UI
flyctl tokens create deploy -a openexec-api-dev -x 999999h
flyctl tokens create deploy -a openexec-ui-dev -x 999999h
# 4. First deploy
gh workflow run "Deploy (dev)" -f target=both
QA bootstraps the same way against the -qa app names (push to the qa branch, or gh workflow run "Deploy (qa)"). See docs/deployment.md for the full runbook (operations, rollback, common failure modes, why .flycast isn't used).
The deployed UI is gated behind Google sign-in with an email allow-list, and the public API is protected by a shared-secret header between the UI proxy and the FastAPI backend. See docs/auth.md for the full setup (Google Cloud Console steps, required Fly secrets, adding/removing users, rotating secrets, and a debugging table).
All settings via environment variables. Minimum required: ANTHROPIC_API_KEY β
unless you configure a local or OpenRouter backend instead (see Running on
Local Models). At least one provider must be set or
the app refuses to start.
| Variable | Required | Default | Description |
|---|---|---|---|
ANTHROPIC_API_KEY |
YesΒΉ | β | Anthropic API key |
DEFAULT_MODEL |
No | claude-sonnet-4-6 |
Executive + most specialists |
DEEP_REASONING_MODEL |
No | claude-opus-4-7 |
CSO, CFO, GC, Board |
VECTOR_STORE_PATH |
No | ./chroma_db |
ChromaDB directory |
EPISODIC_DB_PATH |
No | ./episodic_memory.db |
SQLite for episodic memory |
COMPANY_PROFILE_PATH |
No | ./company/profile.yaml |
Company profile |
ENABLE_CACHING |
No | true |
Anthropic prompt caching |
ROUTING_MODEL |
No | claude-haiku-4-5-20251001 |
Model for intent routing |
SLACK_BOT_TOKEN |
No | β | Slack bot OAuth token |
SLACK_APP_TOKEN |
No | β | Slack socket mode token |
EXEC_EMAIL_ADDRESS |
No | β | Executive Gmail address (Gmail MCP OAuth) |
EMAIL_POLL_INTERVAL_SECONDS |
No | 60 |
How often to poll for new email |
TELEGRAM_BOT_TOKEN |
No | β | Telegram bot token (from @BotFather) |
TELEGRAM_WEBHOOK_SECRET |
No | β | Random string for webhook validation |
DISCORD_BOT_TOKEN |
No | β | Discord bot token (Developer Portal β Bot tab) |
DISCORD_APP_ID |
No | β | Discord application ID (General Information tab) |
DISCORD_GUILD_IDS |
No | β | Comma-separated guild IDs for dev slash-command registration |
DISCORD_NOTIFY_CHANNEL_ID |
No | β | Default channel ID for outbound notifications |
GOOGLE_CHAT_PROJECT_NUMBER |
No | β | GCP project number for Google Chat |
GOOGLE_CHAT_SERVICE_ACCOUNT_FILE |
No | β | Path to service account JSON key |
GOOGLE_OAUTH_CLIENT_ID |
No | β | Google OAuth client ID (Gmail MCP) |
GOOGLE_OAUTH_CLIENT_SECRET |
No | β | Google OAuth client secret (Gmail MCP) |
OPENROUTER_ENABLED |
No | false |
Route Claude calls through OpenRouter and unlock non-Anthropic models per-agent in the Council UI |
OPENROUTER_API_KEY |
No | β | Required when OPENROUTER_ENABLED=true |
LOCAL_MODELS_ENABLED |
No | false |
Route selected slugs to a local OpenAI-compatible server (Ollama, LM Studio, vLLM, llama.cpp) |
LOCAL_BASE_URL |
No | β | Local server URL incl. version path, e.g. http://localhost:11434/v1. Required when LOCAL_MODELS_ENABLED=true |
LOCAL_API_KEY |
No | β | Optional bearer token (vLLM / gateways); Ollama & LM Studio need none |
LOCAL_MODELS |
No | β | Comma-separated local model slugs to surface in the Council UI and route locally, e.g. llama3.3,qwen2.5 |
LOCAL_TIMEOUT_S |
No | 300 |
Per-call timeout for local generation, in seconds |
HONCHO_ENABLED |
No | false |
Per-person memory layer (honcho.dev) β a peer card shared across all channels |
HONCHO_API_KEY |
No | β | Required when HONCHO_ENABLED=true |
HONCHO_BASE_URL |
No | β | Self-hosted Honcho endpoint |
See .env.example for the full list.
ΒΉ
ANTHROPIC_API_KEYis required only when you serve Claude models directly. It can be omitted entirely if you run on local models (LOCAL_MODELS_ENABLED) or route through OpenRouter (OPENROUTER_ENABLED).
Open Executive can run against any OpenAI-compatible local server β Ollama, LM Studio, vLLM, or llama.cpp β instead of (or alongside) the Anthropic API. Local model slugs route to your server through the same provider abstraction the hosted models use; no agent or orchestrator code changes.
# 1. Pull a capable, tool-use-friendly model (example: Ollama)
ollama pull llama3.3
# 2. In .env β point at the local server and list the slugs to expose
LOCAL_MODELS_ENABLED=true
LOCAL_BASE_URL=http://localhost:11434/v1 # Ollama default
LOCAL_MODELS=llama3.3
# 3. (Optional) run with NO Anthropic key β make local the default everywhere
DEFAULT_MODEL=llama3.3
DEEP_REASONING_MODEL=llama3.3
ROUTING_MODEL=llama3.3
# ...and leave ANTHROPIC_API_KEY unset
The listed slugs appear in the Council UI model dropdown, so you can also run a hybrid setup β keep the Executive on Claude while flipping individual specialists to a local model per-agent.
Caveats. Server-side web search (ENABLE_WEB_SEARCH) and Anthropic prompt
caching / extended thinking have no local equivalent and are automatically
disabled for local models. Multi-agent routing leans heavily on tool use, so
pick a model that's strong at it (e.g. Llama 3.3 70B, Qwen2.5) β small models
may route poorly. LOCAL_API_KEY is only needed if your server (vLLM, or a
gateway) requires a bearer token; Ollama and LM Studio need none.
packages/core/openexecutive/agents/your_agent.py extending BaseAgentpackages/core/openexecutive/prompts/domain_prompts.pypackages/core/openexecutive/orchestrator/router.py β add to SPECIALIST_REGISTRY and the specialist enum in SPECIALIST_TOOLSDOMAIN_ALIASES in packages/core/openexecutive/knowledge/retriever.pyknowledge/builtin/your_domain/evals/scenarios/make dev # Start FastAPI + Next.js
make test # Run Python tests
make eval # Run eval suite
make lint # Run ruff + mypy
make docker # Build and run Docker stack
# Unit tests only (no API calls required)
pytest packages/core/tests/unit/ -v
evals/ contains 29 scenarios covering all 8 domains, scored by claude-opus-4-7 as an LLM-as-judge. Each scenario defines a query, simulated company context, expected topics, required specialist routing, and a domain-specific rubric. Five scoring dimensions (persona coherence, domain accuracy, company context utilization, routing quality, actionability) are each rated 1β5. The CI gate requires β₯ 3.5/5 average; any dimension dropping > 10% vs main fails the PR.
Everything in company/ is gitignored β the profile YAML, uploaded documents, and the ChromaDB vector store. None of this leaves your local machine (or your own Fly volume in cloud deployments) except as part of prompts sent to the Anthropic API. Anthropic does not train on API data.
See .github/CONTRIBUTING.md. All PRs must include:
Apache 2.0 β free to use commercially, requires attribution.