Insights
How we built v1 of a new kind of employment marketplace, Part 2: From user stories to a product worth reading
By Saad Ahmed, Founder of Cosmic ·
An agreed structure made it possible to design screens again, but the first screens still read like a pitch deck. The fixes for this employment marketplace v1 were an eight-panel onboarding flow that explains the new model, a rule that processing keeps the full evidence behind each claim, and a cleanup review that shows every change side by side. That design became a v1 built in six days with about 470 automated tests.
Key takeaways
- Judge agent output by whether someone can do the task in it. A polished demo of the concept is a different, easier target.
- Explain a new model in onboarding: four panels per audience, each with one job and one next action.
- Accept the files people already have (PDF, Word, Excel, Markdown, text) and keep normalization, anonymization and publication as separate steps.
- When the product's value is detail, write preservation down as a rule. The agent's default is to summarize.
- Show cleanup changes as editable before-and-after comparisons. A summary of what changed doesn't show what happened to the text.
At the end of Part 1, our team and the founder who hired us had stopped designing screens long enough to agree on the people, their needs and the structure of the product. Candidates would build, refine, approve and share a work profile. Recruiters would discover or receive a profile, evaluate the relevant experience and decide whether to interview.
That structure let us design screens again, but it didn't make them good.
Codex's next version looked like a presentation about an application. It had clean screens, a gallery and explanations of the flows, and it felt like something you'd show in a deck rather than something a person would open to get work done.
I wanted Work history to be a working surface. Candidates should be able to bring in source material, read and fix it, and see what's ready and what still needs attention. The distinction mattered because an agent can produce something polished that answers the wrong question. "Can you demonstrate this concept?" and "Can someone use this?" lead to different interfaces.
We had also reviewed outside design skills and written a skill specific to this client's product. It put our decisions ahead of generic UI advice: start with the task, keep the agreed navigation, use the existing architecture, and make agent access part of the design. Clerk was already the authentication foundation. Open SaaS and other repositories were references, and Mantine with CSS Modules was the UI direction. We didn't need another framework to solve a product problem.
The skill gave future agent sessions a starting point. I still had to tell Codex when a result missed.
One reference did help: the new-task screen in a task-tracking app Cosmic had built. A focused task takes over the screen while the surrounding application steps back. We borrowed that for adding work and creating a role, and later it became the basis for onboarding. We took the interaction and left the rest of that app's look behind.
I also made a basic layout rule explicit: content belongs inside white containers, and grey is the surrounding canvas. Page titles, copy, lists and buttons shouldn't float on grey on one route and sit inside a container on another. It sounds small, but it removes a decision from every screen and gives the application a recognizable structure.
How do you explain a new way to hire?
The bigger correction was that the app wasn't explaining itself.
"No resume required" states a position without explaining how looking for work or hiring changes. A candidate needs to understand why a detailed work account is useful, where that account comes from and who will read it. A recruiter needs a reason to spend time with something much longer than a resume.
We moved that explanation into a short, full-screen onboarding flow: four panels for candidates and four for recruiters. Each panel had one job and a clear next action.
The candidate opening became direct:
The world is changing. Resumes aren't cutting it.
Then we explained what the product makes visible: the problem, the decisions, the person's contribution and what happened next.
The next panel started where people already work. Their meetings, calls and projects leave a record in tools such as Glean, Zoom, Teams and Gong. That record usually serves the employer. Candidates can use the permitted parts of it to reconstruct their own experience. We added recognizable workplace-tool logos, which made the idea easier to grasp than another paragraph about "source material."
I kept pushing on the copy. "Start with the work you already have" sounded harmless but abstract. "Your work is already recorded. Make it work for you" actually said something. We removed helper sentences that repeated the heading or filled space without helping anyone act.
What files can candidates bring?
We also stopped being opinionated about the wrong thing: Markdown.
The original idea had grown around a long career Markdown file. That didn't mean candidates should have to produce Markdown. They should be able to bring PDF, Word, Excel, Markdown or plain text, and the product should handle normalization. A review, project document or meeting record is useful for what it contains, whatever the file type.
The file panel became "Add files that show what you did," with icons for reviews, project notes and meeting records. It explained cleanup and anonymization, followed by the candidate's responsibility to check the draft before sharing. We kept normalization, anonymization and publication as separate concepts. Extracting text successfully doesn't make it safe to publish.
At one point, onboarding had both a workplace-tools panel and a separate panel explaining how to get a workplace prompt. The second repeated the first, so we removed it. Four panels were enough.
The final candidate panel had to explain the payoff. I didn't want to end on administrative instructions about approval. The reason to do this work is that a recruiter and their AI might recognize experience a resume would never reveal. Someone may have solved a difficult operational problem without ever having "operations" in their job title, and the product should give that work a chance to be found.
What should a recruiter see first?
We mirrored the story for recruiters. The promise was that the recruiter, their AI and the product could examine a much deeper account of someone's work together, with the hiring decision left to the recruiter. The profile would open decisions, outcomes, chronology and gaps to scrutiny.
That changed the recruiter's order of operations too. Our first checklist asked them to create a role and connect an agent before they had seen a full profile, which was backwards. A recruiter should first read an example and understand what's different. Once they see its depth, suggesting agent-assisted analysis makes sense, and role creation and wider discovery can follow.
The founder settled on a free one-day trial or $80 a month for recruiter access, with candidate-invited reading kept free. Those became product decisions in the mocks and specification; trial activation and billing weren't built yet. Recruiter accounts also stayed individual: sharing an employer's email domain wouldn't give coworkers access to one another's activity.
Together we also resolved what sharing meant. A candidate's link lasts five days. A recruiter who signs up and saves the profile within that window can keep reading until the candidate revokes access. That created an acquisition path without letting a paid subscription override the candidate's control. Contact stayed lightweight, and recruiter agents would never receive identity or contact details.
After onboarding, we first sent candidates to a separate setup page, which created another destination to return to. I moved those actions into Work history instead: an expanded "Build your profile" section, followed by a separate container for the actual work. Completed actions show green checkmarks without becoming dead ends, so someone who has added files can still add more.
Why does the product keep every detail?
The largest lesson came when we used the founder's own career export to test cleanup.
Codex produced a shorter, rewritten profile from selected sections. It kept some useful examples and numbers but dropped much of the call-by-call evidence. When I asked whether it was a direct copy, it explained the reduction.
That was the wrong instinct for this product.
The product exists to preserve the detail that lets someone check whether a person can do the job. A call can show how someone asks questions, handles objections, separates an assumption from a fact or changes course. Cutting that detail for length removes the evidence.
We made preservation a canonical rule: keep the full work narrative and supporting call evidence; retain deal values, chronology, nuance and uncertainty; and separate career-planning advice and outreach drafts for review instead of silently deleting them. Summaries can help a recruiter navigate, but they can't replace the underlying account.
That rule went into the marketing copy, onboarding, build specification, agent guidance and a machine-readable contract for implementation. The shortened draft stayed as a UI example. The founder's original export stayed intact and became the first real evaluation case for the cleanup engine.
The founder's sales background made another detail obvious. A salesperson might want a client's name replaced with its industry while keeping the deal value. Removing every number would destroy most of the commercial meaning. The system also has to preserve what each number represents, because revenue earned, revenue at risk, renewal upside and customer spend are different claims.
How do candidates review cleanup changes?
Then we had to decide how a candidate reviews the changes.
Our first answer was a document editor with tracked changes and comment bubbles: accept, edit, keep private, undo. It worked as a prototype, but I questioned whether version 1 needed that much machinery. Could we simply show what changed and ask whether it looked right?
The first simplification went too far. A summary saying "we replaced client names" didn't show what happened to the text. We changed it to concrete before-and-after comparisons, then made the replacement column editable so an applied correction updated the matching passages in the profile.
Document-review software gave us a useful reference: persistent hit highlights and a difference viewer. The relevant idea was to keep the document readable, show differences in context and let the reviewer move between them. We didn't need the surrounding review workflow.
The resulting mock combined a readable profile, highlighted changes and a compact way to inspect the original and the replacement. Edits reset cleanup approval, and removed links stay visible as removals. We kept one important boundary: accepting cleanup doesn't publish the profile.
We also removed overlapping controls. "Something needs changing" became redundant once the replacement text was editable. "Add more work or context" served a different need: bringing in material the draft didn't contain yet.
For that, a lightweight writing screen was enough. Candidates can add their own account, use sensible headings and bullets, and pick an approximate period. We chose broad defaults: the last six months, months seven through twelve, and the last one to two years. Work can cross those boundaries. The eventual system can infer dates where the source supports them, leave uncertainty visible and allow correction.
We didn't need to build Google Docs, a page-layout tool or a rigid career timeline to give someone authorship.
What changed for mobile?
Mobile forced another simplification. A full comparison table followed by a long profile is a lot to manage on a phone. We made the profile the main reading surface. Tapping a highlight opens a bottom sheet with the original, the replacement and an Apply change button. The complete list of changes sits behind a control, approval stays in a bottom action bar, and adding more context belongs at the end of the profile.
For onboarding, we standardized the viewport budget instead of adjusting each panel by eye. All eight panels and their calls to action were checked at six desktop, tablet and phone sizes, for 48 passing layout combinations. That verifies the mock at those sizes. It doesn't replace testing accessibility, persistence or real processing in the application.
What did v1 ship with?
Before moving toward implementation, we named five templates: onboarding takeover, guided workspace, task dialog, profile reader and focused task page. Each screen got an identity, content ownership, entry conditions and a next action. We documented where the approved design ended and unbuilt backend behavior began.
That is where this stage landed: onboarding that explains the new model in eight panels, a Work history people act in, and a cleanup review that shows every change.
The work with Codex was a series of corrections: stop presenting and start working; explain the new model; remove the repeated panel; preserve the evidence; show the actual changes; put the recruiter in front of a profile first.
The agent made each revision quickly. Our job was to notice when a convenient shortcut, like the shortened career export, had changed what the product was for.
The design work above became the v1 build: a deployed, tested stack with 30 database tables, row-level security and about 470 automated tests, six days from spec. The case study covers the build and what it's designed to do for the founder's revenue.
Frequently asked questions
Why did the AI agent's cleanup of the career export fail?
Codex shortened the profile and dropped much of the call-by-call evidence, the detail a recruiter needs to judge how someone works. We made preservation a written rule in the spec, the agent guidance and a machine-readable contract.
How do candidates review what the cleanup changed?
They read their profile with each change highlighted in place and open any highlight to compare the original with the replacement. The replacement is editable, edits reset approval, and accepting cleanup never publishes the profile.
How long did the v1 build take?
Six days from spec to a deployed, tested stack with 30 database tables, row-level security and about 470 automated tests. Billing was the remaining piece before launch.