How sharp PMs think, work, and lead when the product they are building can surprise them.
"You are not shipping features.— Jahid Hasan
You are managing the tendency of intelligent systems."
The mindset change nobody warns you about when you move into AI product management.
Your first week. You have just come from a senior PM role where the software did exactly what you told it to. Click a button, it turns blue. Every time. You understood the system because you defined it.
Then your research lead sends a message: "Checkpoint 47 has a 14% regression on our benchmark. But the model's writing quality actually feels significantly better. Do we gate on the test score or trust the feel?"
You write back, "Let me dig in." You open a new tab. You quietly search what "benchmark regression" means in this context. This is where your education actually starts — not when you signed the offer, but here, in the gap between what you knew and where you have landed.
Building an engineering manager platform taught me something strange. That's how I saw it. I found almost every PM, TPM, and EM (somewhat they almost always exist together) who moves into AI product work passes through that gap. Most are good at their jobs. Many have years of shipping behind them. None of it fully prepares them for one simple truth:
In traditional software, you manage behavior. In AI, you manage tendency.
A button turns blue on hover because a developer wrote a rule. A language model has no single answer. It learned patterns from billions of examples, and when you ask it something it produces what it calculates is most likely correct. Not always the same answer. Not always what you expected. Often very good. Sometimes completely wrong and completely confident about it.
In startups, especially, your job shifts fast. You are no longer defining what the system does. You are shaping what it tends to do, and learning to work honestly with the gap between those two things.
"In traditional software, you manage behavior. In AI, you manage tendency. That is not a subtle distinction — it changes everything about how you plan, what you measure, and how you communicate."
In most product work, uncertainty decreases as you ship. In AI, some uncertainty is permanent. The model can behave differently on the same input. Your job is not to eliminate that — it is to characterize it well enough that your team can make confident decisions despite it. "We don't know yet" is a complete sentence. "Here is what we do know and what we are watching" is even better.
An eval suite is a set of tests that measure whether the model is doing what you want. Think of it as your shared definition of "better." Without one, every team conversation about quality becomes a debate of opinions. With a well-built eval, you can catch regressions before users see them and defend decisions with something more than instinct. If you join a team with no eval suite, your first project is to help build one.
PMs who come from traditional software often treat safety as a checklist at the end. In AI product work, that instinct breaks trust with the people who matter most. Safety researchers carry real organizational authority for good reason. The PMs who build the best working relationships with them are the ones who invite them in at the start, not at the deadline.
AI development moves in cycles: form a hypothesis, run an experiment, evaluate the results, decide whether to continue, adjust, or stop. Then do it again. Sprint planning and roadmaps still apply — but they need to be flexible enough to absorb what the experiment actually teaches you, which is rarely exactly what you expected.
One more thing: this shift takes time. Most experienced PMs arrive with a confident, well-practiced way of operating. Some of it does not transfer directly, and the only way through is to stay curious enough to keep learning while still doing the actual work. The curiosity is the job requirement. Everything else follows from that.
How the best teams make decisions without filling up a calendar and why writing is how real thinking happens.
Three months in. Sunday evening. You open your calendar to prep for the week. It is a solid wall of meetings. Research sync. Engineering sync. Safety review. Stakeholder update. Launch readiness check. You have spent nearly half your week in rooms and still feel like you are missing the conversations where things actually get decided.
Then a senior PM shows you her calendar. Nearly empty. You ask how. He/She says: "We stopped meeting to share information. We only meet to debate the document."
Meetings feel like progress — you are all in the same room, talking, engaged. But when a meeting is the only record of a decision, that decision lives only in the heads of the people who were there. Someone misremembers. Someone missed it. Someone joins three weeks later and has no idea it happened.
Written decisions compound. You can reference them, share them, open them at midnight. Writing is also how you actually think hard about something — the act exposes the gaps that conversation covers over.
"A decision document that ends with 'here are three options, you decide' is not a document. It is a meeting agenda with better formatting. The recommendation is the entire point."
One paragraph framing the situation. Two or three genuinely different options, with honest trade-offs for each. One specific recommendation, stated as a sentence, with the reasoning behind it. That last sentence is the document. If you cannot write it,if you find yourself writing "there are merits to both approaches" — you have not finished thinking yet.
Written at the start of any significant workstream, before anyone builds anything. Three questions:
One page. A brief written at the start prevents the most expensive kind of waste: teams building the right thing in the wrong direction.
Sent every Friday. Designed to be read in under five minutes. One line at the top — (red), (yellow), or (green) — with a single sentence explaining why. Three things that happened. Three things next week. One request if something is blocked. The people reading it should know everything they need before Monday, without a single meeting.
Written after something goes wrong. The only rule is that it is blameless. A post-mortem that assigns blame produces an organization that hides failures next time. One that analyzes systems produces an organization that catches the same problem earlier next cycle.
Before committing to something significant, write it up and invite colleagues to comment in writing. Written comments are almost always more technically honest than what people say in a meeting. In a meeting, they want to be supportive. In a document, they think it through first.
Every meeting in a well-run async team has a document — written before the meeting, read by everyone before they arrive. The meeting is not for catching up. It is for debating what the document got wrong. Meetings used for information transfer are a sign the writing isn't happening yet.
What an AI launch demands that no other kind of launch does and the gate framework that keeps chaos from being a surprise.
The marketing video was done. The blog post was ready. The launch readiness deck had every item marked green. Everything was on track.
Then, four days before launch, the trust and safety lead posted in the channel: "We've identified a consistent way to extract system instructions from the model. It works reliably. We need to pause for mitigation."
The 72 hours that followed were chaos not because the problem was impossible to fix, but because nobody had a process in place for exactly this situation. Everyone had assumed a green deck meant they were ready. Nobody had asked what "ready" actually meant.
An AI launch is not a software release with extra steps. Research is finishing one model version while safety is reviewing the previous one. Marketing is building around capabilities that might still change. Legal is reviewing language for edge-case behaviors nobody has fully characterized yet. The PM's job is to make the invisible visible — proactive project visibility that turns scattered work into a picture the whole team can act on, not a deck everyone assumed was green.
In the final six weeks before any major launch, run one short meeting per week — twenty minutes, no more. Everyone has already read the status document. The PM opens the tracker and reads only the items marked red or yellow. For each one, the person responsible explains what is blocking progress. Green items are not discussed. Teams that do this consistently surface blockers four to six weeks before they would otherwise appear. That extra time is the difference between a solvable problem and a launch crisis.
Any launch-blocking problem that shows up less than two weeks before the target date is a process failure, not just a technical one. The gates, the tracker, and the readiness review exist to surface problems earlier. If they didn't, ask why not, and fix it for next time.
At some point you will need to tell people the launch is moving. Be direct, be plain, and come with a new date you can defend. "We found a safety issue that needs more time. New target is the 18th. Here is what changed and why we're confident in the new date." That is the full message. A delay surfaced two weeks out is a logistics problem you can solve together. Two days out, it is a fire and you caused it by waiting.
RAID logs, pre-mortems, and the habit of looking four weeks ahead instead of one.
The engineer appeared at the PM's desk at 4 p.m. with an expression he/she had not seen before. "We have a problem." The latest model checkpoint was producing better outputs, but the compute cost had jumped 40%. At launch scale, the cloud bill would exceed budget by a factor of three.
This had been visible in the research team's internal notes for two weeks. The information was there. It just never made it to the PM, to engineering leadership, or to finance. By the time it surfaced, there were eight days until the planned launch date. What had been a manageable tradeoff became a crisis because nobody had a system for surfacing it earlier.
There is a category of PM capability that rarely appears in job descriptions: the ability to feel a problem before it becomes urgent. It is not a sixth sense. It is structure — a clear understanding of the categories of things that go wrong in AI projects, combined with a habit of asking the right questions on a regular cadence.
"The blocker you find four weeks early costs you a conversation. The one you find four days early costs you the launch."
Four to six weeks before a major launch, run a pre-mortem with your core team. Tell the group: imagine it is one week after launch, and things went meaningfully wrong. Not catastrophically — just badly enough that we are all disappointed. What happened? Everyone writes independently for five minutes, without talking. Then each person shares one item. You go around until the list is empty. Group similar items. Identify the two or three risks that are both likely and high-impact. Assign a clear owner and a mitigation plan to each. Put them on the RAID log.
Teams that value momentum and AI teams do, intensely make it professionally uncomfortable to raise potential failures. The pre-mortem reframes this. Raising a concern is not pessimism. It is due diligence. The format gives people permission to voice what they have been privately worried about. Run it before the launch pressure is real, while there is still time to act.
How to work across teams that barely share a vocabulary and why translation is the PM's most underrated skill.
The kickoff had eight people from six functions. The research scientist was explaining benchmark performance. The safety engineer was asking about coverage for misuse vectors. The policy manager was asking about regional regulatory implications. The designer was trying to understand what "benchmark performance" meant for real user experience. Everyone was smart. Barely anyone understood what anyone else was actually saying.
The PM stopped the meeting. "Let me try to translate," he/she said.
The PM in an AI organization is not the most expert person in the room on any single topic. That is not a gap to close. That is the job. Your value is the ability to hold all the domains simultaneously, understand enough of each to have a real conversation, and bring what you hear back to the people who need to hear it.
"The PM's job in a cross-functional room is not to advocate for a position. It is to translate everyone else's — accurately, without distortion, even when that translation makes the PM's preferred option look worse."
| Function | Their core concern | What they need from you |
|---|---|---|
| Research | Whether the model is actually improving in ways that matter for users | Clear success criteria and protection from scope creep |
| Safety | Failure modes that could cause real harm at scale | Early involvement, not a last-minute sign-off request |
| Engineering | Whether what they're building will still be the right thing in six weeks | Stable priorities and clear answers on ambiguous requirements |
| Design | Whether users will understand what the product can and cannot do | Honest information about model behavior and its limits |
| Legal / Policy | What claims are defensible and what regulatory exposure exists | Enough time to review properly — not a 48-hour request |
Every month, pick one domain you work with and spend two hours learning more about it. Ask a safety engineer to walk you through a red-team session. Sit in on an eval analysis. You will not become an expert. You will become someone who can have a real conversation — which is exactly what the job requires.
Conflict in AI organizations is almost never about personalities. It is about genuinely different beliefs about how much risk is acceptable.
The eval hit 85%. Six weeks ago, that number had felt like an ambitious goal. Engineering was ready to ship. The team had worked hard and earned it.
Then the safety lead sent a note: "The 15% failure rate includes consistent cases where the model gives harmful advice when prompted in a specific way. The pattern is reliable and replicable. We can't sign off at this threshold."
Engineering's position: the eval was the agreement, the eval passed, we ship. Safety's position: the eval was a proxy for safety, not safety itself, and the proxy has a flaw. Both positions were coherent. Both teams were right from within their own frame. The PM stared at the thread for a long time before picking up the phone.
Most conflict in AI product organizations does not come from bad intentions or difficult personalities. It comes from people who care deeply about their part of the problem, each of whom is right within their own frame of reference. The disagreement is not about facts. It is about how much risk each party is willing to carry.
| Type | What it looks like | How to resolve it |
|---|---|---|
| Factual | "The eval passed." / "The eval doesn't capture what we care about." | Get the facts on the table together. Often resolves when everyone sees the same data. |
| Values-based | "Moving fast matters." / "Getting this right matters more." | Escalate explicitly. This is a leadership decision about organizational values, not a team-level call. |
| Process-based | "We should have defined success criteria earlier." | Acknowledge it, decide for now, and fix the process for next time. |
Diagnosing the type correctly is the first move. Most people argue as if it is factual when it is actually values-based. You cannot resolve a values disagreement by adding more data. You can only resolve it by naming it clearly and getting the right people in the room.
When consensus is not achievable and a decision must be made: acknowledge the disagreement explicitly and on the record. Make a decision transparently and name who made it. Document the minority view and the reason it was not chosen. Then ask every team member including those who disagreed to commit to executing the decision as if they had agreed from the start.
The step most teams skip is documenting the minority view. Without it, the same disagreement resurfaces three months later. With it, you have a record. You can show what was considered and why the team landed where it did.
"Unresolved conflict does not go away. It goes underground and surfaces later as slow execution, passive resistance, and a launch that nobody fully owns."
Some PMs avoid escalating because it feels like admitting they should have resolved it themselves. This instinct is wrong. Escalating a genuine values disagreement to the right leader is a correct assessment of which decisions require which levels of authority. Escalate cleanly: state the disagreement clearly, present both positions fairly, and bring a recommendation. You are not asking your manager to solve the problem. You are bringing them the context to make the call.
What executives actually need from you and why most status updates miss the point entirely.
3:17 in the morning. A journalist had tweeted about a hallucination from the beta — a confident, specific, completely wrong answer about a medical topic. The CEO sent four words to the product channel: "What is our exposure?"
The PM sent a link. It pointed to a risk assessment document he/she had written and published internally nine days earlier laying out exactly this failure mode, its frequency, its severity, and the mitigation already in progress. The CEO read it in three minutes. "Okay. Keep me updated." He/She went back to sleep.
The most important principle in executive communication feels counterintuitive to PMs trained to be thorough: synthesis is more valuable than completeness. A six-paragraph update forces the executive to process six paragraphs to find the two sentences that matter. Your job is not to share everything you know. It is to share what they need to act. These are very different tasks.
"No surprises is not a goal. It is a promise. The PM who keeps that promise becomes the person leadership trusts with information early and that is where real influence lives."
The instinct when something goes wrong is to wait for good news to pair with it. This instinct is almost always wrong. Bad news ages very poorly. A problem surfaced three weeks before a deadline is solvable. Three days before, it is a fire.
The formula: state the facts plainly, without emotional framing. Own the problem — not the blame, but the problem. Present two or three paths forward. Recommend one and say why. Ask for alignment on the direction. That is the message. Making it longer usually makes it worse.
For any critical slide or topic, prepare three versions: the full version, a one-paragraph summary, and a single sentence. Different executives enter at different levels of detail. Knowing which version to lead with is part of reading the room something you learn only by doing it a few times and paying attention to what lands.
PM, TPM, and EM look similar from the outside. They are three entirely different functions and they only work when all three know exactly where they end.
At the first standup of the new quarter, three people gave a status update on the same milestone. The Product Manager, the Technical Program Manager, and the Engineering Manager. Each described a slightly different picture. None of the three timelines fully agreed.
An engineer in the back raised her hand: "Which one of them is actually in charge?" None of the three had a clean answer not because they were not capable, but because nobody had ever had the explicit conversation about where each person's authority began and ended.
"The PM asks what. The TPM asks how. The EM asks who. When all three questions get answered well and consistently, teams can do things that look extraordinary from the outside."
| Role | The question they own | Where they spend their attention |
|---|---|---|
| PM | What are we building, and why? | User needs, product direction, trade-offs, stakeholder alignment, success definition |
| TPM | How does it all fit together across teams? | Dependencies, timelines, cross-team coordination, risk tracking |
| EM | Who builds it, and can they do their best work? | Team health, technical quality, hiring, capacity, delivery execution |
In the first week, block an hour with just the three of you. Talk through: Where do our roles overlap, and who owns the call when they do? How will we handle genuine disagreements on direction? How do we want to communicate — when, and in what format? It is an awkward conversation. It is also the most valuable sixty minutes you will spend together all quarter.
What certification gives the AI PM and the honest point where the map runs out.
He/She had passed the PMP exam six months before joining the AI lab. He/She had studied carefully — critical path analysis, stakeholder engagement, risk registers, earned value management. He/She had a binder, color-coded by domain. He/She was proud of it.
On his/her third day, a researcher told him/her that the model's training run the central event around which the entire project plan was organized might need to be restarted due to a data quality issue. The restart would take two weeks. He/She opened his/her binder. There was no section for this.
The project management body of knowledge was built across decades of practice in industries where the work, while complex, is fundamentally predictable. Even Agile assumes that what you are building has discoverable properties. AI development does not fit this model. The model has emergent properties that nobody knows in advance including the people who built it. This does not make the PMP irrelevant. It makes it necessary but not sufficient. Every principle holds. Every specific technique needs to be translated.
| Classic PM Practice | What it becomes in AI product work |
|---|---|
| Stakeholder management | Add safety teams, regulators, and the model's emergent behavior as an implicit stakeholder you cannot fully control |
| Sprint planning | Experiment-driven, eval-gated cycles with no fixed sprint length — the cycle ends when the experiment concludes, not on a calendar date |
| Critical path analysis | Probabilistic planning that treats model checkpoint quality as a variable, not a known input |
| Risk register | A living RAID log where risks change shape — a jailbreak found in red-teaming can flip a launch-blocking item from green to red in the same week |
| Definition of done | An eval suite the whole team agreed on before the work started without it, "done" is whatever the most senior person in the room says it is |
"The PMP teaches you to manage complexity. AI teaches you to manage complexity that changes shape while you are managing it. Both skills are necessary. Neither is sufficient alone."
All of them are learnable. The PM who builds them deliberately — who treats their own professional development as a product to ship, with intentions and a plan — is the one who finds themselves with real authority in AI organizations. Not because they accumulated credentials, but because they became genuinely useful in situations where most people are out of their depth.
That is the through-line of this book. Not a set of frameworks to copy, but a way of thinking about the work — what it requires, where it is genuinely hard, and how to build the judgment to navigate it well. The appendix turns that judgment into an operating system for proactive execution: project health you can see early, cadence that moves work forward, and process lean enough to ship.
You are not shipping features. You are managing the tendency of intelligent systems in a world still figuring out what that means. That work matters. Do it well.
The operating system for proactive execution — project health visibility that moves work forward.
Use this appendix when you need a diagram to pin on a wall or share with a new teammate. The chapters teach judgment; this section is the OS — see problems early, signal health clearly, align before execution, automate the repeat work. RAID itself is covered in Chapter 04; async writing culture in Chapter 02; PM / TPM / EM roles in Chapter 08. New to the team? Start with First 90 Days (A.6). Standing up your stack? See PM Tooling (A.7).
Proactive visibility means surfacing risk and project health before execution stalls — then acting with the lightest process that works. If a ritual does not improve speed or signal, it does not belong in the stack.
Daily async by default — sync only when the work needs depth, alignment, or a decision.
Layer meetings by horizon. The daily pulse stays lightweight; each tier below adds people, time, and strategic weight. Default to writing first; reserve live time for blockers, dependencies, and decisions that cannot close async.
Typically at the start of the workday with the core dev team. Same three questions every time. Prefer an automated prompt (e.g. Slack bot) that collects distributed posts, rolls a digest, and flags anything for the parking-lot after-party — a short sync only for people who need it.
| Cadence | When | Who | Purpose |
|---|---|---|---|
| Daily stand-up | 15 min · daily · start of day | Core dev team | Three questions — yesterday · today · blockers. Live sync at day start, or transition to async (bot prompt · distributed posts · digest) to eliminate the meeting when the team is ready. |
| Weekly sync | 60 min · weekly | Team leads, owners, key contributors | Cross-functional dependencies · tactical adjustments · identify risks · review individual progress · triage immediate priorities. Send agenda beforehand. |
| Bi-weekly program sync | 90 min · every 2 weeks | Program team — PM, relevant EM, design lead | Progress vs. program milestones · holistic program health · major milestone review · risk impact & scope changes · ensure alignment |
| Monthly stakeholder update | 60 min · monthly | Executive leadership · key stakeholders · interested parties | High-level progress · strategic wins · forecasts · critical issues · decisions needing stakeholder input · dashboard of key metrics |
| Quarterly planning | ½ day – full day · quarterly | Leadership · program leads · tech arch · business stakeholders | Define next-quarter OKRs · set strategic priorities · allocate resources · build program roadmaps · shared understanding of what and why |
Don't jump to tickets. Walk the room through a fixed sequence so priorities, capacity, and outcomes stay linked.
Never use a longer meeting to do a shorter meeting's job. Weekly syncs triage — they don't re-plan the quarter. Monthly updates inform — they don't debug tickets. If the daily digest has no blockers, skip the after-party. If quarterly planning ends without written OKRs and a roadmap, it was just a long conversation.
How to keep the log alive after you build it — see Chapter 04 for the framework.
The RAID categories are already in the book. This is the operating rhythm: capture fast, review weekly, escalate when items go stale.
| Letter | Track weekly |
|---|---|
| R | Probability · impact · mitigation owner |
| A | Still true? · what breaks if it isn't |
| I | Severity · next action · target close date |
| D | Predecessor status · who you're waiting on |
A lens for coverage where nothing critical falls through the cracks.
| # | Domain | Ask yourself |
|---|---|---|
| 1 | Plan & Schedule | Do we have a roadmap everyone believes? |
| 2 | Risk & Issues | Is the RAID current and owned? |
| 3 | Comms & Stakeholders | Does each audience get what they need, when they need it? |
| 4 | Meetings & Docs | Are decisions written down and findable? |
| 5 | Budget & Performance | Are we tracking cost and progress honestly? |
| 6 | Team & Resources | Is capacity realistic for the plan? |
| 7 | Learning & Growth | Is the team getting better, not just busier? |
Who does the work, who owns the call, and who needs a seat at the table.
Before a major milestone, map each decision and deliverable to four roles. RACI removes the "I thought you had it" conversations that stall AI launches.
| Letter | Role | Ask |
|---|---|---|
| R | Responsible | Who is executing? |
| A | Accountable | Who answers if this fails? (exactly one) |
| C | Consulted | Who must weigh in before key decisions? |
| I | Informed | Who should be kept in the loop on progress or outcomes? |
A starter matrix for a cross-functional launch. Adjust roles to your org; keep the single-Accountable rule.
| Activity | PM | TPM | EM | Research | Safety |
|---|---|---|---|---|---|
| Product scope & trade-offs | A | C | C | C | I |
| Cross-team timeline | C | A | R | C | I |
| Engineering delivery | C | C | A | I | I |
| Eval suite & model quality | C | I | I | A | C |
| Safety sign-off | C | C | I | C | A |
| Launch comms & stakeholders | A | R | I | I | I |
One A per row — if everyone is accountable, nobody is. Every A needs at least one R. Use C sparingly; too many consulted parties is just another meeting in disguise. Revisit the matrix when the team or milestone changes.
How autonomous teams, clear ownership, and customer-value metrics turn execution into revenue and retention.
PM work only matters when it moves customer outcomes and business results. These four ideas connect how teams run day to day to what leadership actually measures.
| Concept | Business question it answers |
|---|---|
| North Star Metric | Are we building what drives customer value and revenue? |
| One source of truth | Is everyone deciding from the same facts? |
| Ownership | When this slips, who answers and who executes? |
| Self-managing teams at scale | Can teams move fast without losing alignment? |
The metric that best reflects value customers receive. When teams argue about priorities, it reframes the debate: does this move the number that predicts retention and growth?
Decisions scattered across threads, decks, and side conversations create rework and slow growth. One living document — roadmap, RAID, RACI, launch tracker — is where the org looks before debating.
Growth stalls when ownership is fuzzy. One person is Accountable for the outcome; others are Responsible for execution. Issues get solved when names are on the row, not in the hallway. See RACI (A.4).
High-performing teams decide and deliver without constant direction. As you add teams, representatives sync in a Scrum of Scrums — dependencies and blockers surface before they hit customers or revenue.
When growth slows or launches slip, ask in order: Are we measuring the right North Star? Is there one source of truth? Is ownership named on the RACI? Are self-managing teams coordinated before blockers become customer-visible? That sequence is the proactive execution loop — fix the system before adding more people.
Learn the system, build visibility, then optimize — without reorganizing everything on day one.
A playbook for PMs and TPMs joining a team (or founders standing up one). Each phase compounds on the last. Pair with RAID, RACI, and async stand-ups.
Listen more than you fix. Your job is to build an accurate picture of how work actually flows — not how the onboarding deck says it flows.
| Focus | Actions |
|---|---|
| Structure | Map PM / TPM / EM lines · who decides what · upstream & downstream teams |
| Workflows | Shadow stand-ups, planning, launches · note where information dies |
| Tracking | Stand up RAID log · one source of truth for status |
| Quick wins | Pick 1–2 low-risk fixes that build trust — not a reorg |
Every week, publish the same written format. Use Red / Amber / Green as project health at the top — one signal, then the detail. This is the heartbeat of the proactive execution OS: health you can read in ten seconds, detail for those who need to act. Pull shipped work from your tracker (e.g. Linear) when it keeps wins honest.
A green RAG with hidden blockers is worse than no update. The point is to surface drift before the timeline slips — then assign owners and move. If the update does not change what someone does that week, cut it.
🟢 RAG status: Green / Amber / Red — project health
Overall progress (1–2 sentences): where we are vs. plan this week.
Top 3 wins — bullets, things shipped or unblocked this week.
Top 2 blockers — active issues delaying the timeline · who owns the fix.
Next week's focus — absolute highest-priority milestones.
Dependencies are a planning problem, not a surprise you discover mid-sprint. Manage them in weekly planning, not during execution.
| Practice | What good looks like |
|---|---|
| Define | What exactly is needed · from whom · by when |
| Own | One DRI (directly responsible individual) per dependency on the RAID log |
| Assess | Blockers and risks logged before the week starts — not after slip |
| Sync | Short regular check-in with each DRI · review status in weekly planning |
The goal is to close coordination gaps, not add process weight. Skip meeting-heavy status forums, excessive approval chains, and long planning cycles that slow engineering velocity. Prefer small incremental improvements — and flag openly when a new ritual adds complexity without speed.
Turn what you learned into rhythm: stakeholders know what to expect, metrics show trends, and the RAG update runs every week without a meeting.
One improvement per cycle. Automate what repeats. Remove what adds friction.
| Theme | Incremental moves |
|---|---|
| Process | Refine sprint planning · streamline reviews · better docs |
| Dependencies | Weekly planning owns the map · DRIs synced · blockers removed before execution |
| Efficiency | Async-first · automate reminders & broadcasts · fewer status meetings |
| Performance | KPIs tied to outcomes · recognize wins in the RAG update |
| Distributed teams | Clear handoffs · reliable cadence · explicit ownership — like a well-run distributed system |
Transparent decisions in async environments need a written frame — context, options, owner.
| Framework | Use when | Roles |
|---|---|---|
| RACI | Ongoing deliverables & ownership | Responsible · Accountable · Consulted · Informed |
| DACI | Single decision with a deadline | Driver · Approver · Contributor · Informed |
| RFC | Significant change before commitment | Problem · options · pros/cons · timeline · recommendation |
Big goals fail when they stay abstract. Break them down the way reliable systems break down load.
What did we learn about the team? What broke that we should have seen earlier? What one process change do we try next month? Publish answers in writing. Founders and small teams need this discipline as much as large orgs — transparency scales down, not just up.
A practical stack — plan, write, build, measure, automate. Pick one winner per layer, then wire them together.
Tools don't replace judgment. They remove friction so you spend time on decisions, not copy-paste. This is the automate layer of the proactive execution OS — wire the stack so visibility and project health update themselves. The examples below reflect what modern AI product teams commonly run; your exact vendors may differ, the layers shouldn't.
| Layer | Common tools | What to automate | Why AI PMs need it |
|---|---|---|---|
| Plan & track | Linear · Jira · Asana | Status sync · sprint boards · dependency links · shipped-this-week for RAG updates | Single source of truth for what ships. Linear is common at fast-moving AI startups; Jira still dominates large enterprise program offices. |
| Write & align | Notion · Confluence · Google Docs · Slack | PRD templates · RFC comments · async stand-up digests · decision logs | AI work needs written context — prompts, eval criteria, launch gates. If it isn't documented, it didn't happen. |
| AI build & eval | OpenAI API · Anthropic API · Google Gemini / Vertex · ChatGPT · Claude | Prototype flows · draft specs · red-team prompts · run eval suites before launch | Model providers the industry actually ships on. PMs don't need to train models — they need API access and a repeatable eval loop. |
| Eval & observability | LangSmith · LangFuse · Weights & Biases · custom eval harnesses | Regression tests on prompts · trace failures · compare model versions | What OpenAI, Anthropic, and top labs internalize: quality is a product surface. Track it like uptime. |
| Measure & learn | Amplitude · Mixpanel · Looker · Metabase · Datadog | Usage funnels · feature adoption · cost per request · executive dashboards | Connect model behavior to customer and business outcomes — not just "the demo worked." |
| Automate | Slack workflows · Linear automations · Zapier · n8n · cron + webhooks | Daily stand-up prompts · RAG broadcast · stale-issue nudges · RAID reminders | Replace the coordination work you keep doing manually. This is where cadence from A.1 actually runs itself. |
Exact stacks vary by team size and compliance — but patterns repeat. Use this as a benchmark, not a shopping list.
| Org type | Typical stack signals | PM focus |
|---|---|---|
| AI labs (OpenAI, Anthropic, etc.) | Own models + APIs · heavy internal eval tooling · strict launch review · Slack + docs culture | Eval gates, safety sign-off, written launch criteria — speed with guardrails |
| Big tech product (Google, Meta, Microsoft, etc.) | Jira or internal trackers · Confluence/Notion · Figma · enterprise analytics · multiple model vendors | Cross-org dependencies, executive dashboards, quarterly OKR tooling |
| AI startup | Linear · Notion · Slack · OpenAI or Anthropic API · lightweight analytics · Zapier-style glue | Move fast: one tracker, one doc hub, automate stand-ups and status on day one |
Buy for the workflow, not the logo. One tracker, one doc system, one chat hub, one analytics source — then automate the handoffs. Model access (OpenAI, Anthropic, or your org's approved vendor) is non-negotiable for AI PMs; everything else should earn its seat by removing a meeting or a spreadsheet. If a tool adds coordination overhead, cut it.
Jahid Hasan
First-Principles AI Product Leader · Entrepreneur
Jahid Hasan is an AI-native product builder who co-founded and led product development at Tometo AI. The company grew out of what he saw at Microsoft — large teams stalling not on engineering, but on coordination. Traditional PM process added visibility theater without faster execution.
Tometo AI gives engineering teams proactive project visibility so they can move faster with less friction. His work favors lean process and clear ownership: optimize what accelerates delivery, cut what only adds system overload.