Pheebs is an open-source AI telemetry tool created by Eversynced that measures how engineers and teams actually work with AI coding agents. It sits quietly inside Claude Code, Cursor, and Codex via hooks, capturing lightweight interaction signals: the shape of the session, not its contents. The people it is built for are the ones who need an honest proficiency read rather than a guess — engineering leaders, platform teams, and the developers themselves. Its purpose is measurement: turning the ordinary activity of agent sessions into signals about model choice, verification habits, context management, orchestration, and the dollars that model choices are costing.
The problem Pheebs addresses is a visibility gap that opens up precisely when a team starts moving fast. AI coding agents arrive, adoption climbs, and nobody can say what changed. Six observations illustrate the questions the tool was built to answer: model spend that buys nothing, such as a bigger model than the work needed; where AI code ships unchallenged; whether AI output gets verified at all; rework hiding inside the speedup, where follow-up prompts are fixing something the AI broke; and whether the enablement investment landed — for example, a review skill used weekly by 78% of engineers while a migration skill never caught on. The final observation frames the stakes: nobody on the team runs tests inside the agent loop, which is a missing harness rather than a skills gap. Distinguishing a structural gap from a coaching gap is the core problem Pheebs exists to solve.
Pheebs captures data by hooking into the agents themselves. It works with three coding agents — Claude Code, Cursor, and Codex — and records seventeen event types that run from session_started through to artifact_found. Hooks fire on session starts and ends, prompt submissions, skill and slash-command expansions, sub-agent spawns, tool calls and failures, compaction, and background tasks. Typical recorded fields are deliberately small: a session_started event carries a codebase such as acme/checkout and a model such as opus; a prompt_submitted event carries a character count and, when the prompt intent classifier is enabled, an intent label such as task or debug; a tool_use_completed event carries the tool name and its duration, with recognized commands summarized as a tool_intent such as test_run. Claude Code and Codex additionally export native OpenTelemetry metrics and logs through the Pheebs proxy, while Cursor is covered by hooks alone.
The design constraint behind all of this is that Pheebs captures interaction patterns, not content. It never records source code or file contents. It never records file paths or directory structures — a repo is reduced to org/repo from the git remote. It never records prompt text; a prompt becomes a character count. It never records raw command strings, since a command like npm test is read in process and recorded as tool_intent: test_run. It never stores your name or your email: the developer is the id behind your Pheebs token, stamped by the backend, or a truncated hash of your git email when no token is set, and your GitHub handle is never looked up. The only route in the backend contract that receives raw text at all is POST /classify-prompt, which takes one prompt in and returns one label out. The backend contract also includes POST /ingest for one event envelope per request, POST /validate-token to resolve a token to an identity and its consent flags, POST /otel/v1/{signal} as an OTLP passthrough so no observability credential ever ships in the client, and an optional GET /insights for what one developer can see about their own work.
On top of those signals sits a documented proficiency model. It assesses six competencies: Models, covering model choice, effort settings, plan mode, and autonomy modes; Artifacts, the reusable configuration that shapes the agent, such as skills, sub-agents, slash commands, and context files; MCP, live connections to external systems like tickets, databases, browsers, and documentation; Evals, verification wired into the agent loop through tests, typecheck, lint, build, and review passes; Context management, deliberate use of the context window including compaction and the save, resume, clear lifecycle; and Orchestration, running more than one agent at a time via sub-agents, parallel work, worktrees, hooks, and plugins. Each practice is classified as Unobserved, Adopted, or Recurring — Recurring meaning it showed up in at least 3 of the last 4 active weeks — and the coverage index summarizes, per engineer, the share of applicable practices at Recurring.
Five judgement signals sit alongside the competency model. On the output side, verification coverage measures the share of AI edits followed by a verification action such as a test run, typecheck, lint, build, or a check against a spec; pushback rate measures how often the engineer challenges AI output instead of accepting it; the refinement-to-repair ratio separates follow-up prompts that refine intent from those that repair breakage; and wholesale-accept rate captures sessions with no pushback, no repair, and no verification, weighted by lines changed — described as the composite red flag of polished output with no questions asked. On the input side, model-fit rate measures the share of sessions whose model class matched the size of the work. Pheebs follows three stated principles here: tasks are sized, so every task prompt gets a scope from a one-file change to open-ended design and a session is judged on its hardest prompt; misses count both ways, because an over-provisioned session burns budget silently while an under-powered one shows up as repair prompts; and Pheebs is an audit, not a router — it never intercepts a prompt or switches a model on anyone's behalf, it reads the gap and prices it, and the decision stays yours.
The overall pipeline has five steps. A hook fires. Lightweight fields are extracted — event type, durations, counts, models, trigger types — with prompt text reduced to a character count and an optional intent label. Identity and repo are resolved from the Pheebs token or a truncated git email hash, and from org/repo on the git remote. Every event is logged locally in a JSONL log, and with a token set it is also sent to the backend. OpenTelemetry rides along for Claude Code and Codex. Sending requires both settings to be present: pheebs config set base-url and pheebs config set token. Both need to be set or nothing is posted, and unsetting either one stops sending — the local JSONL stays the durable copy either way. Pheebs also states there are four routes to any backend: self-hosted, or managed by Eversynced.
Two deployment shapes are described. In the self-hosted model you run the backend and hold the data: telemetry goes from developers' machines to your infrastructure and Eversynced never sees it. That option includes the full client for all three agents under Apache-2.0, a documented contract and a reference backend in the repo, raw JSONL you can query with whatever you already use, and no account, no key, and no requests. In the managed model, the same open-source client points at a backend Eversynced operates, with the proficiency model rendered as reports and dashboards — the AI Enablement Assessment, a 30-day telemetry sprint ending in an executive debrief and a plan for the gaps. The benefits the content states are visibility rather than surveillance: knowing which models are in play, whether AI output gets verified, whether enablement investments landed, and what the model-fit gap costs.
The reporting built on top includes a practice adoption funnel, with one bar per competency split by how many engineers have not acted on it, acted on it once, or acted on it week after week; a practice heatmap putting every engineer against every competency, where a cold column means the team is missing the setup and practice for it and a cold row calls for coaching; a per-engineer view showing how much of each competency has become habit; and a signals table showing the five judgement signals per engineer against a team median. The dollar view prices the gap: in the illustrative example, a savings opportunity of $9,960 against $32,400 of list-price spend, described as 31% and an API list-price equivalent estimated upper bound, with models used and work as sized split across Frontier, Large, Medium, and Small classes. Decisions and figures come from complete sessions only, with coverage reported as complete, incomplete, no telemetry, and unpriced.
The quickstart is three commands: npm install -g pheebs, pheebs init for interactive setup across all three agents, and pheebs doctor to check the wiring. Eversynced states that every Eversynced engineer is instrumented with Pheebs; it powers the measurement layer of their AI delivery framework, and the reporting built on top of it ships with the AI Enablement Assessment run for client teams. The product is therefore aimed at teams adopting AI coding agents who want evidence about how those agents are actually being used in their codebase.
In short, Pheebs turns the day-to-day shape of AI coding sessions — model choices, prompts reduced to counts, tool calls, verification, compaction, and orchestration — into an honest, legible read on proficiency, adoption, and cost, while keeping the code, the prompts, and the identity of the developer out of the dataset.