The application layer for the agent you already have
The workbench your agent writes into.
A workspace for research that has to be checkable. Your agent does the finding. Everything it finds lands somewhere structured, gets checked against the page it came from by a pass that did not produce it, and is still there next month.
We never call a language model. Not once, anywhere. That is not a cost decision. It is the reason every other sentence on this page can be true.
- 0models we run
- Yours does the thinking, on your machine, with your key. There is no usage meter here and nothing that gets quietly cheaper when prices move.
- 4kinds of work
- Research that has to hold up looks the same whether the subject is a professor, a candidate or a company. One of the four is built and in use today; the others are not.
- 1line to paste
- Your agent already knows how to make this connection. Setting it up takes about a minute, and your key stays where it is.
The draft checker, running on your machine.
That box is the product’s own draft checker, running here, on your machine. Nothing is sent anywhere. Try it on something you actually wrote.
- Vocabulary that reads as generated×4
"delve"
"testament to"
"underscores"
"a testament"
- Praise that labels instead of describing
"groundbreaking"
- A feeling about the work instead of a thought about it
"aligns well with"
- An opener that says nothing
"I hope this email finds you well"
- A closer that asks for nothing×2
"I look forward to hearing from you"
"Thank you for your time and consideration"
- Nothing in it that is specific to this person
103 words with no date, number, venue, title or proper name in them. Every sentence here could be sent to any professor in the field, which is what makes it read as generated. Name the paper, the year, the table, the company -- one real detail is worth the whole paragraph.
7 more findings, not shown here. The app shows every one, against the span of text that produced it.
These are the base rules. A workspace can add its own, and a draft is scanned again on save against the policy in force then. Nothing here is a claim about whether a model wrote it. The finding is that the text reads as generated, which is what the person receiving it will think either way.
The research is not the hard part any more.
Anyone can get an agent to read forty faculty pages. What nobody has is the part after. The transcript is gone by Thursday. The claim that sounded right was never checked against the page it came from. The same person gets researched twice because nothing remembered the first time.
Then the draft goes out reading like every other draft that went out that week, to someone who has learned to recognise exactly that. And when a person asks where a number came from, there is no answer. The work gets done, then it evaporates, and the parts that survive cannot be defended.
In the chat
- What the agent found
- Where it came from
- That you already did this person
- Why you ranked them first
In a workspace
- What the agent founda record
- Where it came froma source and a date
- That you already did this persona duplicate check
- Why you ranked them firstan audit trail
What happens when the platform never calls a model
Most people hear this as a way to save money. The money is the least interesting part of it.
- 01
The free tier is real
A user costs us a database row. Every tool that runs the model for you pays per enrichment, which is why their free tier is a trial.
- 02
Your prompts never reach us
Not as a policy. There is nowhere here to put them. The agent runs on your machine and writes results over MCP.
- 03
No model risk
Nothing to deprecate, no rate limit of ours, and no quiet substitution of a cheaper model to protect a margin. That last one is what usually kills these products.
- 04
You are always on the better model
You are paying a frontier tier already. Anything embedded here would be whatever was cheap enough to embed.
There is one deliberate exception. In hosted mode a paying user supplies their own provider key and the agent loop runs here. We still never pay for a token, and we still never invoke a model on our own account.
A claim is not a fact until something else has checked it.
An agent that checks its own work agrees with itself. So research and verification are separate passes, run by separate invocations, and a finding stays a candidate until the second one has been over it against the page that was cited.
What survives carries the source, how close that page was to the fact, and the date it was read. Staleness is answerable because the timestamp was recorded at the time rather than inferred later.
Where two sources disagree, it stays a disagreement. A product whose whole argument is verification cannot round uncertainty down to confidence anywhere.
Your agent, over MCP
- ›search_professors{ "query": "representation learning", "university": "…" }← 0 records
- ›upsert_professor{ "title": "Associate Professor, tenure track", "source": "cs.…/people/chen" }
- ›upsert_professor{ "accepting_students": "Yes, Fall 2027", "source": "chenlab.…/join" }
- ›upsert_professor{ "group_size": 6, "source": "chenlab.…/people" }
- ›upsert_professor{ "funding": "NSF grant, renewed" }
- ›verify_professor{ "mode": "independent" }
- ›score_match{ "against": "applicant_profile" }
- ›create_draft{ "strategy": "formal" }
The record it writes into
Empty.
Nothing on file. The workspace is empty, so the agent goes and reads.
How it fits together
- 01
Connect your agent
One line of MCP configuration in Claude, Cursor, ChatGPT or anything else that speaks the protocol. Your key never leaves your machine.
- 02
It writes into the workspace
Findings arrive as structured records rather than prose. Sources, tiers and read dates come with them, because the tools will not accept a claim without one.
- 03
You work in the UI or you do not
Everything the agent can do is on a screen too, and everything on a screen is a tool. Neither is the second-class path.
~/.claude/mcp.json, or Settings → Connectors
{
"mcpServers": {
"flint": {
"type": "http",
"url": "https://flintworkbench.com/api/mcp",
"headers": { "Authorization": "Bearer <your token>" }
}
}
}Four readers, one habit
Four readers, one habit: find out about someone, check what you found, write to them, and be able to show your working a month later. One of these is built and in use. The other three are described here and not written, and this page will keep saying so until they are.
Graduate applicants
Sixty professors. One workspace. Every claim checked.
You are writing to sixty professors in the season you are also finishing a thesis. The ones who reply are the ones who could tell you had read something. Import the roster, have your agent research it, and write outreach that could only have been sent to one person.
The hard questions
Do I need a separate account for each product?
No. One login covers the family, and adding a product to it does not mean signing up again. Workspaces are separate though, one per product, because a candidate pipeline and a professor roster are not the same shape, and the one place that tried to hold both would serve neither. Your grad workspace and your recruit workspace sit side by side under the same account and never see each other.
Can I not just build this myself?
Yes, and some people will. I did: seventy-eight thousand lines over two and a half months. The code was the easy half. The hard half is knowing that verification has to be a separate pass or the model agrees with itself, that a source needs a reliability tier and a read timestamp or staleness becomes unanswerable, and that a faculty roster saying “affiliate” regularly means a tenure-track professor. That knowledge only accumulates by getting it wrong first.
Why not use Claude or ChatGPT on their own?
You should, and you will. Bring them here. What a chat cannot do is remember across sessions, keep findings structured, notice it researched this person in July, prove where a claim came from, or tell you which of sixty people to write to first. The transcript is gone by Thursday. This is where it lands instead.
What am I paying for, if you do not do the AI?
The filing cabinet, the fact-checker and the checklist. Somewhere findings live, a pass that checks them against the page they came from, and a procedure that is the same this week as it was last week.
I already have a spreadsheet.
A spreadsheet has no provenance, no freshness and no verification, and your agent cannot write to it in a way that stays structured. It also will not tell you the person in row 34 changed institution in March. Spreadsheets are genuinely fast though, which is why the lists here are sortable, dense, editable in place and support bulk actions.
Setting up MCP sounds technical.
It is one line to paste and takes about a minute. For some people that is still too technical, and the honest answer there is hosted mode, where the agent loop runs here using a provider key you supply.
Is anyone paying for this?
No. One person uses it, the author, for a real application cycle, which is how every real bug in it was found.