The Research Agent That Replaced My First Hire: Setup, Prompts, and 4 Hours a Week
William DeCourcy · August 24, 2026
For most small operations, the first hire on the whiteboard is somebody to do the looking-things-up. Pull the background before a call, check what the competition changed this month, read the last 40 replies and say what people keep asking. The work that has to happen before the real work starts.
I planned that hire too, and then I didn't make it. Professor Leads runs as a team of one, and the job I would have hired for is the one I handed to a research agent instead.
A research agent is a standing job with a defined output, a schedule, and a verification step, and that structure is the whole difference between owning one and having a habit of asking a chatbot good questions. The habit produces nothing on the weeks you forget. The job produces the same artifact whether you remember or not.
Key Takeaways
- Gathering is the job worth automating first. A research agent assembles the raw material for content, competitors, and customers, and every call about what it means stays with you.
- 3 properties turn a chatbot habit into an agent: a named output artifact, a recurring calendar slot, and a verification step. Miss any one and you're back to asking questions when you happen to think of it.
- The 3 assignments that pay first are content prep (roughly 2 hours a week), competitive intel (about 1), and customer research (about 1). That's the 4 hours, and it's gathering time rather than thinking time.
- Charge the honest tax: verification runs 45 to 60 minutes a week on 3 running assignments, so the real recovery lands nearer 3 hours. A tool that claims zero verification cost is charging it somewhere you can't see.
- The prompt structure that works names the sources, fixes the output shape, and requires a citation per claim. Open-ended prompts return essays; constrained prompts return artifacts you can act on.
- The trust problem is different from searching your own notes. Your archive can't invent a source, and the open web can, so citation accuracy is where these agents fail most and where your check belongs.
The hire you were about to make
Think about the job description you'd actually write for that first hire. Most of it is retrieval: find the thing, summarize the thing, put it where I can see it before the meeting.
I've sat through enough hiring conversations at AmeriLife and before it to know how that role gets justified, and the justification is almost always time rather than judgment. The pitch is "I'm spending 6 hours a week on prep I shouldn't be doing," and it's usually true.
What that means in practice is that the first hire is a time trade dressed up as a capability. What you're actually buying is the return of hours you're spending on work that doesn't need you specifically.
That's exactly the shape of work an AI agent covers well, and it's why this particular hire is the one worth reconsidering before you make it. The parts of the job that need a person (reading a room, deciding what matters, owning a relationship) are the parts you weren't going to delegate anyway.
What makes it an agent and not a habit
Most people who say they use AI for research mean they open a chat window when a question comes up. That's useful, and it's also why the time savings never show up in a calendar.
What's missing? Structure. 3 properties turn the same underlying tool into something that behaves like a hire.
A named output artifact. The agent produces a specific document that lands in a specific place: a one-page brief, a change log, a ranked question list. If you can't name the file, you don't have an assignment yet, you have a topic.
A recurring slot. The work happens Monday at 8 whether or not you thought of it. An agent that runs when you remember is a habit with better branding.
A verification step. Before the output gets used, somebody checks the sources. This is the property people skip, and it's the one that decides whether the whole thing is an asset or a liability.
Simple.
The reason the structure matters more than the model choice is that the structure is what compounds. Week 6 of a running assignment is better than week 1 because the agent has context and you've tuned the prompt, and none of that accumulates if you're starting from a blank chat window every time.
The 3 assignments worth handing over
I run 3, and I'd start any operation with the same 3 in this order. Each one has a clear artifact and a clear owner for the decision that follows.
Content prep
The assignment: before I write anything substantial, assemble what's already been said about it, what the current numbers are, and where the consensus is soft.
The artifact is a one-page brief with 5 sections: the claim most people make, the evidence behind it, the best counter-argument, the freshest data point with its source and date, and 3 questions the existing coverage leaves open. That last section is where the piece I actually write usually comes from.
This is the assignment that saves the most time, because content prep is the work that expands to fill whatever afternoon you give it. It's also the one where a bad output is cheapest to catch, since I'm going to read every source anyway before I build an argument on it.
Competitive intel
The assignment: watch a fixed list of competitors and tell me what changed.
The artifact is a change log. Pricing page edits, new positioning language, a feature announcement, a shift in who they're clearly talking to. Each entry gets a date, a link, and one line on what it might mean.
Why a change log and not a summary? Because a summary of a competitor is the same every week, and the thing you need is the delta. A weekly report on 6 competitors takes 20 minutes to read and teaches you nothing; a change log with 3 entries takes 2 minutes and occasionally changes what you do.
Keep the list short. 5 or 6 companies whose moves would genuinely change something you do.
Customer and lead research
The assignment: before a call, tell me who this is and what they probably need. Across the month, tell me what the people writing in keep asking.
There are 2 artifacts here and they run on different clocks. The per-call one is a 5-line brief: the company, the role, the likely problem, the 1 thing they've published or said that's relevant, and the question I should open with. The monthly one is a ranked list of the questions that keep coming up across replies, forms, and calls.
That second artifact is the sleeper. A ranked question list built from your own inbound is the best content calendar input available, and it rarely gets built because building it by hand is miserable.
The setup, in 3 moves
You can stand all 3 assignments up in an afternoon. The setup is the same for each.
1. Give it sources you trust. Point the agent at a defined set: the competitor URLs, your own past writing, the 4 or 5 publications you'd actually cite, your call notes if you keep them. An agent pointed at the whole internet returns answers from internet results, which is the thing you were trying to improve on.
2. Give it a shape. Write the output format into the prompt, section by section, and be specific about length. "A brief" gets you an essay; "5 sections, 3 bullets each, 1 line per bullet, every factual claim followed by its source URL and date" gets you something you can use.
3. Give it a slot. Put each assignment on the calendar at a real time, and put the verification 15 minutes after it. Monday 8:00 the content brief runs, Monday 8:15 I check it. The check has to be scheduled or it becomes optional, and an unverified brief is worse than no brief.
The prompts
The prompt is where these live or die, and the pattern is the same across all 3 assignments. Name the sources, fix the shape, require a citation per claim, and tell it what to do when it doesn't know.
Here's the content prep prompt, close to what I actually run:
You are preparing a research brief for a piece I'm writing on {TOPIC}.
Sources: only use material published in the last 18 months. Prefer
primary sources (original research, filings, platform documentation,
first-party data) over summaries of them.
Produce exactly these 5 sections:
1. THE CONSENSUS CLAIM: what most coverage asserts. 2 sentences.
2. THE EVIDENCE: what that claim rests on. 3 bullets, each with a
source URL and publication date.
3. THE BEST COUNTER-ARGUMENT: the strongest case against it, stated
fairly. 3 sentences.
4. FRESHEST DATA POINT: the most recent relevant number, with source
URL and date. If nothing published in the last 90 days, say so.
5. OPEN QUESTIONS: 3 questions the existing coverage does not answer.
Rules: every factual claim gets a source URL and a date. If you can't
find a source for a claim, drop the claim and note what's missing.
Do not summarize the sources back to me; tell me what they establish.
3 lines in that prompt do most of the work. The recency window stops it from handing me a 2019 blog post as current, and "tell me what they establish" is what keeps section 3 from turning into a book report.
The one worth stealing outright is "if you can't find a source, drop the claim." It converts a confident guess into a visible gap, which is exactly the behavior you want from something you're going to trust.
The competitive intel prompt swaps the sections for a change log and adds one rule that matters more than the rest:
Compare each URL against your previous run. Report ONLY what changed.
If nothing changed for a competitor, write "no change" and move on.
Never pad the log to make the week look eventful.
That last line exists because I've watched these tools invent significance on a quiet week. An agent that reports "no change" 4 weeks running is doing its job correctly, and you have to say so explicitly or it'll find something.
The 4 hours, worked
Here's the arithmetic, because a time-savings claim deserves to be shown rather than asserted.
Content prep ran about 2 hours a week for me: reading around a topic, chasing down current numbers, finding the counter-argument I should address. Competitive intel ran about 1 hour, mostly clicking through the same 6 sites and trying to remember what they looked like last month. Customer research ran about 1 more between pre-call prep and the occasional attempt to pattern-match my inbox.
4 hours. That's the number in the headline and it's the gathering time, which is genuinely what moves.
Now the honest part. Verification isn't free, and anyone quoting you a savings number without it is quoting you a gross figure and calling it net.
On 3 running assignments I spend 45 to 60 minutes a week checking sources: opening the cited links, confirming the dates, and spot-reading the ones a claim actually rests on. So the real recovery is closer to 3 hours, and that's the number I'd plan against.
Is 3 hours worth an afternoon of setup? For a team of one, that's about 150 hours a year of work you were doing that didn't need you. It pays for the afternoon in the first fortnight and keeps paying.
The better argument is where the 3 hours land. They come back as time in front of the decision instead of the gathering, and the decisions are the part of the job that's actually yours.
The trust problem here is different
I've written before about building a search agent on top of your own notes, and the trust question there is mostly about retrieval: did it find the right thing in your archive, and can it show you where.
The research agent has a harder problem. Your own notes can be wrong, but they can't invent a source that never existed. The open web can, and models are demonstrably worse at citations than at almost anything else they do.
Published benchmarks put citation accuracy at the bottom of the pile, with reported hallucinated-reference rates running from the low teens well into the majority depending on the model and the subject. Treat that as an argument for building the check in from day 1 rather than bolting it on after something embarrassing ships.
What does that check look like in practice? 3 things, and it takes about 15 minutes.
Open every cited URL, all of them. The characteristic failure is a plausible-looking URL that 404s or lands on a page that never said the thing, and a spot-check of 2 out of 8 will sail right past it.
Confirm the date on anything time-sensitive. A correct fact with a 3-year-old date is a wrong answer when the claim is about what's happening now.
Read the source for any claim you're about to build on. If a number is going into your writing, your pricing, or your pitch, you read the page it came from. Everything else can ride on the summary.
That's the routine, and after a few weeks it stops feeling like overhead. You develop a fast sense of which assignments run clean and which ones need a closer look, and the check compresses to the claims that carry weight.
Where this goes wrong
Running it without a verification slot. This is the failure that turns a useful agent into a liability, and it happens gradually. The first 3 weeks check out clean, the check starts feeling like a formality, and the first hallucinated citation you ship is one you'd have caught in 40 seconds.
Handing it the decision. The agent tells you a competitor changed their pricing page. It should not tell you what to do about it, and if you find yourself asking it to, you've quietly moved the job from gathering to deciding.
Pointing it at everything. An agent with no source list returns the same average-of-the-internet answer your prospects can get themselves. The source list is most of the value, and narrowing it is usually the fix when output quality drops.
Building 6 assignments at once. Same failure as building 9 email segments. Stand up 1 assignment, run it for 3 weeks, and add the next only when the first one is producing something you actually read.
Judging it on the first run. Week 1 output is the worst it will ever be, because the prompt is untuned and the agent has no context for what you find useful. Give an assignment a month before you decide whether it earns its slot.
The trade
You give up the version of this where a person absorbs the ambiguity. A real hire figures out that you didn't mean the whole pricing page, just the enterprise tier, and adjusts without being told. An agent needs that written down, and writing it down is the work.
What you get back is an assignment that runs on a quiet week, doesn't need managing, and produces the same artifact in month 9 that it produced in month 1. For an operation of 1, consistency is worth more than intuition.
And you get the hours back in the right place. Gathering has always been the tax on getting to the job, and this is the first tool I've run that reduces the tax rather than moving it around.
The hire I didn't make is still the right hire eventually. When it happens, that person walks into an operation where the prep is already assembled and their first week can be spent on something harder than looking things up.
Further Reading
On Professor Leads
- Your Second Brain Needs a Search Agent covers the inward-facing half of this: an agent that answers from your own notes, calls, and documents, plus the 5-question test for when to trust it.
- The 5-Tool AI Stack for a One-Person Growth Team is where the Researcher role first got named as 1 of the 5 jobs worth a slot, and it covers the handoff problem that decides whether a stack works at all.
- The Lead Quality Audit is the scoring companion to the customer-research assignment above. Once the agent tells you who a lead is, the audit gives you a repeatable way to decide whether they're worth your morning.
On Forbes (by William DeCourcy)
- The Symbiotic Future: Where Human And Machine Intelligence Meet makes the broader case this post applies to one job: the split worth drawing is between the work a machine does tirelessly and the work that needs a person's judgment.
- Cultivating Excellence: Building A High-Performing Team is the counterweight, and it's worth reading alongside this one. Deciding which roles an agent covers is the same exercise as deciding what you're really hiring a person to do.
William DeCourcy
William DeCourcy is the founder of Professor Leads, Founding and Immediate Past President of the Insurance Marketing Coalition, and a Forbes Business Development Council contributor. He's spent 15+ years in performance marketing, leading teams at Marriott Vacations Worldwide and AmeriLife (where he became the world's first Chief Lead Generation Officer), and built Professor Leads to teach what actually works.

