How to Produce World-Class Research with Perplexity Computer & Claude -> Fable 5
The full methodology and workflow behind a 22-page, 12-chart research report: Perplexity Computer for data collection and essential research, then Opus 4.8 and Fable 5 for verification and production.
The full methodology and workflow behind a 22-page, 12-chart research report: Perplexity Computer for data collection and essential research, then Claude on Opus 4.8 and Fable 5 for verification and production, along with my elite research skill file that holds the whole thing to an institutional standard. The skill’s architecture is broken down below, and the workbooks for building your own are attached.
Too Long →Didn’t Read
Perplexity Computer produced a 22-page research report and a full data workbook in 33 minutes across roughly 140 agent tasks. It still contained 5 substantive errors.
The core doctrine: never publish research from inside the machine that produced it. Move the output to a second, different model and audit it against a written standard.
The audit runs on a skill file with 6 sections: a closed 3-tier source architecture, 4 research modes, search and paywall discipline, 3 output templates, voice rules, and 8 quality checks.
Charts get rebuilt from the raw data in your own visual identity. The final report is written in your own voice and read by you in full before anything ships.
Attached to this post: the full report this article walks through, plus the Your Research Employee workbooks in 4 editions (Claude, ChatGPT, Gemini, Perplexity) so you can build your own analyst in about an hour.
The report runs 22 pages.
It carries 12 original charts, 4 parts, 2 appendices, and a complete source register with a live URL behind every figure.
The machine research behind it took 33 minutes across roughly 140 tasks.
The verification pass that followed caught 5 errors before a single chart was drawn.
The complete process took my team of agents 65 minutes. The work would have taken me two weeks.
Anyone can get an AI research report in 2026. The tools are extraordinary and improving quarter over quarter. Which means the differentiation has moved. World-class research is no longer about access to information. It is about the discipline you wrap around the machines and agentic teams you deploy to make it happen.
This piece walks through the exact pipeline I used to produce “The Last Bubble Went Public. This One Is Staying Private,” my quantitative history comparing the dot-com bubble of 1995 to 2003 against the AI boom of 2022 to 2026. I recorded the whole build on Loom. This is the written version, plus the part the video could not show: the full anatomy of the elite research skill that runs quality control, and the workbooks for building a research employee of your own on Claude, ChatGPT, Gemini, or Perplexity.
Hope this helps you this week, friends!
- j -
If you want to become AI-fluent this summer and learn how to build a team of agents to do the work discussed in this article and so much more, you should seriously consider joining my Oper(AI)tors Summer of AI Fluency cohort.
Use code TAKE500 to receive a $500 discount on your seat.
What is the best AI stack for producing research?
Perplexity Computer collects. Claude verifies and produces the output in my primary stack. Skills encode the standards so neither step depends on my memory or my mood.
Three tools, three jobs, no overlap.
The research ran in Perplexity’s Comet browser (download today!), which is my daily browser on every machine I own. The verification, chart production, and final write-up ran in Claude Cowork sessions on Opus 4.8 and the Fable 5 model. Dictation ran through Wispr Flow, which is why I can commission a research project at the speed of speech.
Step 1: How do you brief an AI research agent like Perplexity Computer?
I opened Perplexity Computer and dictated a brief. Not a query. A brief. The distinction decides the quality of everything downstream.
Watch the Loom here.
The brief had 4 components:
The history: Deep research on the dot-com bubble and crash, from the late 1990s through the early 2000s. I want to understand the period, not skim it.
The comparison: Set that period against the current AI cycle, quantitatively and qualitatively. Startups then against startups now. Trajectories then against trajectories now.
The data demand: Come back with financial data and historical data sufficient to build charts, graphs, and visualizations. This clause is the one most people skip. If you do not ask for chartable series, you get prose, and prose cannot be re-plotted in your own brand.
The standard: Exhaustive, attributable, accurate. Every claim is sourced.
Then I left my computer for a bit, and Perplexity deployed a team of agents that worked for 33 minutes across roughly 140 discrete tasks. Watching the full run is instructive, partly because Perplexity shows which models it selects for which activities. But you do not need to watch. You need the 2 deliverables it returns: a complete quantitative and qualitative report, and an Excel workbook holding every underlying data series as chartable tables.
View the Perplexity Excel data file here.
That workbook is the asset. The report is a draft. Hold that hierarchy in your head, because the next step depends on it.
How I Built a $200,000 Newsletter Business From a Substack Nobody Was Reading
The Five Moves That Turned My Substack Into a $200,000 Business
The Operating Project: Most Coaching Ends With Notes. This Ends With a Machine.
Step 2: How do you verify AI research for accuracy?
Here is the doctrine at the center of this whole system. Never publish research from inside the machine that produced it.
Perplexity Computer is a serious research engine. The output is attributable, the writing is clean, the coverage is broad. And it still shipped me 5 errors. Not typos. Substantive, number-level errors of the kind that end up in someone’s investor memo and then in someone’s screenshot.
So the first thing I do with any Perplexity research package is take it out of Perplexity. I downloaded both files, opened a Claude Cowork session, attached them, and ran my elite research skill against the full package with one instruction: check that everything here is accurate, and show me what is not.
The audit came back (read the full audit MD file here) with a list.
The 5 that mattered most
Cisco’s crash was understated by 22 percentage points: The figure that circulates for Cisco’s dot-com decline is 66 percent. That number measures only through December 2000, which is 22 months before the actual bottom. The true peak-to-trough is 88 percent, from March 2000 to October 2002, from $79 to $9.50. The distinction changes the lesson. Being right about the technology protected Cisco’s business. It did not protect the price.
Anthropic’s valuation ladder was mislabeled: The verified sequence reads $61.5 billion, $183 billion, roughly $350 billion, then $965 billion, with the $965 billion Series H dated May 28, 2026. Versions dating that figure to November 2025 conflate it with the Microsoft-Nvidia deal, valued at roughly $350 billion. A 15x repricing in 14 months deserves correct dates.
The AI venture funding series stopped a year early: The draft ended the series at 2024, at $114 billion. The full record runs through 2025 at $202.3 billion, close to half of all venture capital deployed worldwide that year. Comparisons ending at 2024 understate the cycle. Flattering the past is the most common failure mode in bubble analysis, and it usually arrives disguised as a data cutoff.
The marquee AI IPO count was stale: The correct count is 2. Cerebras was listed on the Nasdaq on May 14, 2026, joining CoreWeave’s March 2025 listing.
The NASDAQ year-end close needed pinning: The verified 2025 monthly series puts it at 23,241.99.
I read the list, approved all 5 corrections, and instructed that all work going forward be updated against them. Every downstream chart and paragraph automatically inherited the corrected numbers. The final report states each correction openly in its methodology appendix, which is what a reader deserves and almost no AI-assisted research bothers to do.
The general principle: collect everything in one place, then verify with a different model, ideally a stronger one than the one that did the collection. Two models will not share the same blind spots. One model grading its own homework will.
This is the first layer of checking, not the last. When the full report is assembled, I read the entire thing myself. The machine audit removes the errors a human skims past. The human reader catches what no rule can encode. You need both, in that order.
Step 3: How do you rebrand AI-generated charts in your own visual identity?
The research you publish should look like yours. Perplexity’s charts are competent and generic. So the next instruction to Claude: take my personal brand skill, take every data series from the workbook, and re-create every chart plus any additional visualizations the data supports.
Claude returned 12 charts. Each one is rendered in my visual identity, the Forest Ink field register, as an individual PNG, plus a single PDF holding all 12, plus an HTML index page for reviewing the full set in one scroll. NASDAQ’s 6.2x five-year multiple. The casualty table shows that the survivors still fell to 88 percent. The private-marks ladder no public index can show you. The 280-versus-2 IPO chart encapsulates the entire thesis in a single image.
View all twelve charts here.
The proportion and consistency across the set is the point. A reader cannot articulate why a report feels institutional. Uniform chart grammar is most of the answer.
Two practical notes. First, this only works because the data came back as chartable tables in Step 1. Second, keep the individual PNGs. You will reuse single charts in newsletters, decks, and posts for months. The report is one container for the charts, not the only one.
Step 4: How do you turn AI research into a publishable report?
Final instruction: produce the complete research report, including the charts, written in the Operating voice. Clearly written, professionally written, minimal emotion. Tell the story, walk the reader through the numbers, let the two periods of history do the arguing. Return it as a branded PDF.
It came back as a 22-page document: executive summary, 4 parts, a methodology appendix listing the corrections, and a full source register. Some charts embedded in the first PDF pass imperfectly. This does not bother me. I had Claude produce a Markdown version I could format and embed cleanly myself. Expect this. The last 5 percent of production is still manual, and pretending otherwise is how polished-looking errors ship.
Then the human read. The whole report, my eyes, before anything publishes. The machines compress 2 weeks of analyst work into an afternoon. They do not remove the afternoon.
What goes inside an elite research skill for Claude?
Everything above depends on one file. The elite research skill is a document I have built and tweaked over roughly 6 months, and it is the difference between asking Claude to “check this research” and running an institutional-grade audit with repeatable standards. A skill, mechanically, is just a structured instruction file the model loads when the task calls for it. The leverage is in what the file encodes.
Mine has 6 sections. Here is what each one does and why it exists.
1. What sources should an AI research skill be allowed to cite?
The skill’s foundation is a closed list of approved sources, organized in 3 tiers. Tier 1 is elite publications: the New York Times, the Financial Times, the Wall Street Journal, the Economist, Harvard Business Review, MIT Technology Review, and MIT Sloan Management Review, each listed with its domain and its specific strength. Tier 2 comprises academic research centers: Harvard, Stanford, Princeton, Yale, Chicago Booth, and LSE, each with named search targets, such as specific working paper series. Tier 3 is institutional investment research: Goldman Sachs, Morgan Stanley, Bridgewater, Blackstone, AQR, BlackRock, JP Morgan, plus the top VC firms, a16z, Sequoia, Benchmark, Bessemer, and their peers, with a rule to pull published essays and market maps rather than portfolio announcements.
The operative word is exclusively. The skill instructs the model to cite nothing outside the architecture unless I explicitly ask. This single constraint kills the SEO-farm citation problem before it starts. Most AI research quality problems are source-selection problems wearing a disguise.
Every finding then carries 1 of 4 credibility tags: [PEER-REVIEWED] for Tier 2 academic work, [INSTITUTIONAL] for Tier 3 research, [EDITORIAL] for Tier 1 reporting, and [OPINION] for op-eds and commentary, which must still come from approved sources. I can scan a brief and weigh every claim by the tag beside it.
2. How does the skill match research depth to the task?
The skill defines 4 modes, each with trigger phrases so the model selects the right one from how I naturally talk, each with its own numbered method.
Thesis Research activates on phrases like “I’m writing about” or “I want to argue that.” Its method: restate the thesis in 1 sentence, find supporting evidence, then deliberately hunt complicating evidence, then historical parallels, then 2 to 3 quotable data points, then any disagreements between sources. The output emphasis is narrative ammunition for an article.
Market Intelligence activates on client names and phrases like “competitive landscape.” It works through recent Tier 1 coverage, institutional research, relevant academic work, then risks and opportunities not yet in mainstream coverage. Output emphasis: decision-ready intelligence for advisory work.
Trend Scanning activates on “what are the smartest people saying about” and similar. It restricts to the last 30 to 90 days, then maps consensus, divergence, and the narrative that is forming but has not crystallized. Output emphasis: compressed, newsletter-ready signal.
Deep Dive activates on “comprehensive research” or “everything on.” Broad search, findings organized by sub-theme, a timeline where the topic has one, the key debates mapped, the gaps named, and a meta-narrative synthesized at the end. This is the 10-plus-source mode, and it is the one the dot-com project ran in.
The modes matter because “research this” is not one task. It is at least 4 tasks with different depths, different recency requirements, and different outputs. Naming them lets the file route the work.
3. How should AI research handle search queries and paywalled sources?
The skill specifies exact search patterns, site: operators per source, working-paper queries for the academic tier, and firm-name patterns for institutional research. It sets a floor of 5 searches per request, more for Deep Dive. It mandates recency within 12 months unless the topic is historical.
It also contains the section I consider non-negotiable: paywall discipline. Most Tier 1 sources are paywalled. The skill instructs the model to attempt the full fetch, and when blocked, to extract only what is actually visible, provide the direct URL for me to open through my subscriptions, and tag the finding [PAYWALLED — link provided]. The controlling rule is on one line: never assume content is behind a paywall. Only report what you can actually see. Fabricated paraphrases of paywalled articles are among the most common and least detectable AI research failures. This section is the immune response.
4. What output formats should AI research produce?
Three fully specified templates, chosen by destination. Format A, the Research Brief, feeds articles: thesis, key findings with tags and full citations, complicating evidence, historical parallels, quotable data points, source disagreements, narrative hooks, and a list of paywalled sources to review. Format B, the Advisory Intelligence Memo, feeds client work: executive summary, market context, a key-intelligence table where every row pairs a finding with its source, credibility tag, and a “so what” implication, then competitive landscape, risks and blind spots, and recommended discussion points for the client call. Format C, the Signal Pack, feeds the newsletter: the signal in 1 paragraph, top insights at 2 sentences each, where the smart money disagrees, the emerging narrative, and a prioritized reading list.
Templates sound bureaucratic until you run without them. Unstructured research output is a wall of prose you excavate. Structured output is a component you install.
5. Why should AI research output never hedge?
The skill enforces the Operating voice on the research output itself. Declarative sentences, no hedging, active voice only, findings stated as facts. “Goldman’s research shows X,” never “Goldman’s research seems to suggest X.” A banned-word list that includes “perhaps,” “arguably,” and “it seems.” The reasoning: a research brief is not a neutral academic document. It is intelligence, written with an operator’s eye for what matters. Hedged findings defer the judgment work to future me, and future me is busy.
6. How do you quality-check AI research before delivering it?
The skill closes with 8 verification items the model must clear before delivering anything: every finding cites a specific approved source, every finding carries a tag, nothing off-architecture slipped in, paywalled items are flagged with URLs, the format matches the destination, the voice complies, source disagreements are surfaced rather than smoothed over, and the output is actionable rather than merely informational.
That 7th check deserves a sentence. Models are consensus machines. Left alone, they average conflicting sources into a smooth, confident, wrong middle. Forcing disagreements to the surface is where the most interesting angles live. Goldman versus Bridgewater is an article. Their average is nothing.
The skill also declares its integrations, it feeds my meeting-intelligence pipeline, hands off to my brand system for visuals, and supplies the source architecture when my contrarian filter needs counter-evidence. Skills compound when they reference each other. A lone skill is a tool. A connected set is a staff.
How do you build your own AI research employee?
You do not need 6 months to get a working version of this. With my Oper(AI)te co-instructor, Wessal Khader, I have built the whole construction process into a set of workbooks called Your Research Employee, attached to this post in 4 editions:
Each workbook is a file that an AI can run. You upload it, paste 1 kickoff prompt, and the AI walks you through the build as a guided interview, 1 question at a time. No fields to fill in by hand. The build takes about an hour across 4 steps:
The Research Standard, about 20 minutes: Nine questions that force specificity: the decisions your research actually feeds, the output shape you need, when a fast answer suffices versus a deep dig, which sources you trust on sight, which are noise to you, your domain, and whether you want a recommendation or evidence laid out for your own call. The test of a good standard is that it names a real decision. A brief with no decision behind it is a Wikipedia page.
The Source Rules, about 15 minutes: The side nobody teaches. The workbook ships with a measured floor built from published research on how AI research fails, including the Columbia Tow Center finding that 8 AI search engines got the source wrong more than 60 percent of the time across 1,600 queries. On top of that floor sit 9 banned research habits, uncited claims, single-sourcing, stale data stated as current, marketing repeated as fact, and 5 more, each paired with the positive rule that replaces it.
Sources and Assembly, about 15 minutes: Your trusted-source list, your internal documents, and your question workflow get tested and assembled into 1 paste-ready analyst profile: standard, rules, sources, mode, output format.
The Source Audit, about 10 minutes: A repeatable protocol for scoring any brief against a fixed table. 0 to 1 violations are decision-grade. 2 to 4 means fix the claims and update your rules tonight. 5 or more is a guess, wearing a citation. Each workbook closes with a swipe file of 6 prompts for daily operation, including my favorite, the test drive: make the analyst research a question you already know the answer to, then grade it against reality.
The 4 editions exist because the 4 platforms do this job differently. Claude is the only one that wires your private data in as sources through connectors, and it thinks the longest before answering. ChatGPT produces the strongest long-form synthesis and is the most confident when wrong. Gemini brings the biggest context window and the tightest Google Docs handoff. Perplexity tested most accurately on citations and holds a standing persona most loosely. No edition is best at everything. Pick the edition for the machine you live in, and if you run several, the steps are the same on each.
These are the files we will be using in Oper(AI)te this Summer, starting this Tuesday, the 21st of July
The Claude Edition: The only edition that connects your private data as sources through connectors.
The ChatGPT Edition: Built around Deep Research and Memory.
The Gemini Edition: Built around Deep Research with Google grounding and Workspace.
The Perplexity Edition: Built around Research mode, Labs, and Spaces.
What should you do with this workflow this week?
The pipeline, compressed: commission research with a real brief that demands chartable data. Move the output to a second, different model and audit it against a written standard before you trust a number. Rebuild the visuals in your own identity. Produce the final document in your own voice, then read every word yourself before it ships.
This week, run 1 piece of it. Take any research you are currently relying on, drop it into a different model from the one that produced it, and ask for an audit against the workbook’s source rules. Count the violations. The number will tell you whether you have research or a guess with a citation.
And if you want to build the full system live, this is exactly what we teach in Oper(AI)te, the Summer of AI Fluency cohort I am running with Wessal Khader. Six live builds: a research analyst, a voice employee, a meetings employee, a visual brand, wired into the tools you already use. We start Tuesday, July 21. Through the deadline, code TAKE500 takes $500 off, $1,999 instead of $2,499. Details at unfazedfounder.com/operaite.
The tools will keep improving without your permission. The standards will not.
Frequently Asked Questions
Can you trust AI research from Perplexity without verifying it?
No. Perplexity Computer is the strongest collection engine I use; every claim arrives attributed, and it still shipped 5 substantive errors in this single report, including a 22-point understatement of Cisco’s crash and a mislabeled valuation ladder for Anthropic. Attribution is not accuracy. Audit everything that feeds a decision or a publication.
Why should you verify AI research with a different model than the one that produced it?
Two models do not share the same blind spots. A model auditing its own output tends to re-approve its own errors, because the same weights that produced the mistake evaluate it. Collect in one place, then verify in another, ideally with a stronger model than did the collection.
Which AI is best for research: Claude, ChatGPT, Gemini, or Perplexity?
They do different jobs. Claude is the only one that reads your private data as sources through connectors, and it thinks the longest before answering. ChatGPT produces the strongest long-form synthesis and is the most confident when wrong. Gemini brings the largest context window and the cleanest Google Workspace handoff. Perplexity tested most accurately on citations. The staff you want runs each job on the machine that does it best.
What is a Claude skill?
A skill is a structured instruction file that the model loads when a task calls for it. Mine encodes an approved source list, 4 research modes, search and paywall rules, 3 output templates, voice rules, and 8 quality checks. The file replaces my memory and my mood as the quality-control layer.
What are the most common errors in AI research?
Wrong citations delivered confidently. The Columbia Tow Center found 8 AI search engines got the source wrong more than 60 percent of the time across 1,600 queries. After that: fabricated links, stale data stated in present tense, single-sourced claims, and vendor marketing repeated as fact. Every one of these is testable, which is why the workbooks ship with a source audit.
How long does this workflow take?
The machine research took 33 minutes. The full pipeline, audit, charts, and write-up included, fit in an afternoon, plus the time it takes you to read the final report in full, which is not optional. Building your own research employee from the attached workbooks takes about 1 hour.
Do you need to know how to code to build this?
No. Every piece of this system is a document and a set of prompts. You upload a workbook, paste 1 prompt, and answer questions in plain language. The most technical act in the entire pipeline is downloading a file from one tool and attaching it to another.
About John
John Brewton documents the history and future of operating companies at Operating by John Brewton. He is a graduate of Harvard University and began his career as a PhD student in economics at the University of Chicago. After selling his family’s B2B industrial distribution company in 2021, he has been helping business owners, founders, and investors optimize their operations ever since. He is a career consultant and business operator, asking the question: What is the future of companies?









It's intriguing to consider how this approach could transform various fields beyond economics. Which industries do you think could benefit the most from this structured AI research methodology?
The no-hedging voice rule is interesting to me. Once you ban "it seems," a wrong Cisco number reads exactly like a right one, so the whole system leans even harder on the audit step actually catching things. Maybe that's fine since you read every word before publishing anyway.