The application layer for the agent you already have

The workbench your agent writes into.

A workspace for research that has to be checkable. Your agent does the finding. Everything it finds lands somewhere structured, gets checked against the page it came from by a pass that did not produce it, and is still there next month.

We never call a language model. Not once, anywhere. That is not a cost decision. It is the reason every other sentence on this page can be true.

0models we run
Yours does the thinking, on your machine, with your key. There is no usage meter here and nothing that gets quietly cheaper when prices move.
4kinds of work
Research that has to hold up looks the same whether the subject is a professor, a candidate or a company. One of the four is built and in use today; the others are not.
1line to paste
Your agent already knows how to make this connection. Setting it up takes about a minute, and your key stays where it is.
01Try it

The draft checker, running on your machine.

That box is the product’s own draft checker, running here, on your machine. Nothing is sent anywhere. Try it on something you actually wrote.

Draft checkerruns in your browser
13 blocking4 warnings
  • Vocabulary that reads as generated×4

    "delve"

    "testament to"

    "underscores"

    "a testament"

  • Praise that labels instead of describing

    "groundbreaking"

  • A feeling about the work instead of a thought about it

    "aligns well with"

  • An opener that says nothing

    "I hope this email finds you well"

  • A closer that asks for nothing×2

    "I look forward to hearing from you"

    "Thank you for your time and consideration"

  • Nothing in it that is specific to this person

    103 words with no date, number, venue, title or proper name in them. Every sentence here could be sent to any professor in the field, which is what makes it read as generated. Name the paper, the year, the table, the company -- one real detail is worth the whole paragraph.

7 more findings, not shown here. The app shows every one, against the span of text that produced it.

These are the base rules. A workspace can add its own, and a draft is scanned again on save against the policy in force then. Nothing here is a claim about whether a model wrote it. The finding is that the text reads as generated, which is what the person receiving it will think either way.

02The problem

The research is not the hard part any more.

Anyone can get an agent to read forty faculty pages. What nobody has is the part after. The transcript is gone by Thursday. The claim that sounded right was never checked against the page it came from. The same person gets researched twice because nothing remembered the first time.

Then the draft goes out reading like every other draft that went out that week, to someone who has learned to recognise exactly that. And when a person asks where a number came from, there is no answer. The work gets done, then it evaporates, and the parts that survive cannot be defended.

the day you did the research
day 07d14d21d28d42d

In the chat

  • What the agent found
  • Where it came from
  • That you already did this person
  • Why you ranked them first

In a workspace

  • What the agent founda record
  • Where it came froma source and a date
  • That you already did this persona duplicate check
  • Why you ranked them firstan audit trail
03The bet

What happens when the platform never calls a model

Most people hear this as a way to save money. The money is the least interesting part of it.

YOUR MACHINEFLINTWORKBENCHYour agentClaude · Cursor · ChatGPT · anythingYour model provideryour key, your tokens, your billnever leavesyour machineMCPstructured writesToolsa claim without a source is refusedWorkspacePostgres · candidate and trusted, kept apartVerificationa second pass, reading the cited pagepromotedno path exists
There is no arrow from this side of the drawing to a model, and that is the product rather than a diagram convention. It is why the free tier costs us a database row, why your prompts cannot reach us, and why nothing here degrades the month a provider changes its prices.
  1. 01

    The free tier is real

    A user costs us a database row. Every tool that runs the model for you pays per enrichment, which is why their free tier is a trial.

  2. 02

    Your prompts never reach us

    Not as a policy. There is nowhere here to put them. The agent runs on your machine and writes results over MCP.

  3. 03

    No model risk

    Nothing to deprecate, no rate limit of ours, and no quiet substitution of a cheaper model to protect a margin. That last one is what usually kills these products.

  4. 04

    You are always on the better model

    You are paying a frontier tier already. Anything embedded here would be whatever was cheap enough to embed.

There is one deliberate exception. In hosted mode a paying user supplies their own provider key and the agent loop runs here. We still never pay for a token, and we still never invoke a model on our own account.

04Provenance

A claim is not a fact until something else has checked it.

An agent that checks its own work agrees with itself. So research and verification are separate passes, run by separate invocations, and a finding stays a candidate until the second one has been over it against the page that was cited.

What survives carries the source, how close that page was to the fact, and the date it was read. Staleness is answerable because the timestamp was recorded at the time rather than inferred later.

Where two sources disagree, it stays a disagreement. A product whose whole argument is verification cannot round uncertainty down to confidence anywhere.

One session, played backillustrative

Your agent, over MCP

  1. ›search_professors{ "query": "representation learning", "university": "…" }← 0 records
  2. ›upsert_professor{ "title": "Associate Professor, tenure track", "source": "cs.…/people/chen" }
  3. ›upsert_professor{ "accepting_students": "Yes, Fall 2027", "source": "chenlab.…/join" }
  4. ›upsert_professor{ "group_size": 6, "source": "chenlab.…/people" }
  5. ›upsert_professor{ "funding": "NSF grant, renewed" }
  6. ›verify_professor{ "mode": "independent" }
  7. ›score_match{ "against": "applicant_profile" }
  8. ›create_draft{ "strategy": "formal" }

The record it writes into

Empty.

Nothing on file. The workspace is empty, so the agent goes and reads.

05How it fits

How it fits together

  1. 01

    Connect your agent

    One line of MCP configuration in Claude, Cursor, ChatGPT or anything else that speaks the protocol. Your key never leaves your machine.

  2. 02

    It writes into the workspace

    Findings arrive as structured records rather than prose. Sources, tiers and read dates come with them, because the tools will not accept a claim without one.

  3. 03

    You work in the UI or you do not

    Everything the agent can do is on a screen too, and everything on a screen is a tool. Neither is the second-class path.

~/.claude/mcp.json, or Settings → Connectors

{
  "mcpServers": {
    "flint": {
      "type": "http",
      "url": "https://flintworkbench.com/api/mcp",
      "headers": { "Authorization": "Bearer <your token>" }
    }
  }
}
06Products

Four readers, one habit

Four readers, one habit: find out about someone, check what you found, write to them, and be able to show your working a month later. One of these is built and in use. The other three are described here and not written, and this page will keep saying so until they are.

Graduate applicants

Sixty professors. One workspace. Every claim checked.

You are writing to sixty professors in the season you are also finishing a thesis. The ones who reply are the ones who could tell you had read something. Import the roster, have your agent research it, and write outreach that could only have been sent to one person.

07Questions

The hard questions

Do I need a separate account for each product?

No. One login covers the family, and adding a product to it does not mean signing up again. Workspaces are separate though, one per product, because a candidate pipeline and a professor roster are not the same shape, and the one place that tried to hold both would serve neither. Your grad workspace and your recruit workspace sit side by side under the same account and never see each other.

Can I not just build this myself?

Yes, and some people will. I did: seventy-eight thousand lines over two and a half months. The code was the easy half. The hard half is knowing that verification has to be a separate pass or the model agrees with itself, that a source needs a reliability tier and a read timestamp or staleness becomes unanswerable, and that a faculty roster saying “affiliate” regularly means a tenure-track professor. That knowledge only accumulates by getting it wrong first.

Why not use Claude or ChatGPT on their own?

You should, and you will. Bring them here. What a chat cannot do is remember across sessions, keep findings structured, notice it researched this person in July, prove where a claim came from, or tell you which of sixty people to write to first. The transcript is gone by Thursday. This is where it lands instead.

What am I paying for, if you do not do the AI?

The filing cabinet, the fact-checker and the checklist. Somewhere findings live, a pass that checks them against the page they came from, and a procedure that is the same this week as it was last week.

I already have a spreadsheet.

A spreadsheet has no provenance, no freshness and no verification, and your agent cannot write to it in a way that stays structured. It also will not tell you the person in row 34 changed institution in March. Spreadsheets are genuinely fast though, which is why the lists here are sortable, dense, editable in place and support bulk actions.

Setting up MCP sounds technical.

It is one line to paste and takes about a minute. For some people that is still too technical, and the honest answer there is hosted mode, where the agent loop runs here using a provider key you supply.

Is anyone paying for this?

No. One person uses it, the author, for a real application cycle, which is how every real bug in it was found.