How Jev works
We have implemented TypeSafe's Jev in our Coworkies platform to evaluate coworking space profiles, job ads and blog posts. Each content type has clear scoring criteria. Jev gives us signals about which records meet those criteria and which need attention. We then use Gemini to create improvement proposals.
That experience has made us interested in a wider question: where could this approach help the people running coworking spaces?
TypeSafe calls Jev a System One model, built for fast, focused judgments inside software. The name draws on the idea of fast, intuitive thinking popularized by Daniel Kahneman. Jev reads the context an application supplies and returns structured evaluations that the application can act on. TypeSafe's System One introduction explains the approach.
There are two parts to a request. First, the context: a listing, job description, support message or account record. Second, the questions and criteria: what you want evaluated and what each possible answer means.
Jev supports three question types:
- Choice: select from categories you define, such as billing, facilities or sales, with probabilities for the options.
- Score: assess something against levels you describe, such as how clearly a job ad explains the responsibilities.
- Noul: return a probability between zero and one for a yes-or-no question, such as whether a message requests a cancellation.
Several questions can evaluate the same record independently in one request. The application can then sort the results, apply thresholds or send selected records for further work. TypeSafe documents the details in its question types reference.
The criteria deserve care. For example, a job-ad clarity score could distinguish between a vague role description, some useful responsibilities and a clear account of the work. That gives the score a meaning the team can review. A general instruction to judge whether an ad is "good" leaves too much unspecified.
Why Jev works well for evaluations
Repeated evaluations need clear criteria and results that software can use consistently. Five features make Jev a useful fit for this work.
No free-form text. Each answer arrives as a defined value the application can store, compare and sort. A content review queue can use the score directly, without extracting a verdict from a written response.
Calibrated probabilities and confidence signals. TypeSafe trains Jev so its probabilities reflect uncertainty across many evaluations. Choice and Score answers also include a confidence value summarizing how concentrated those probabilities are. That helps a team route ambiguous records for review. Thresholds still need checking against your own examples. TypeSafe explains confidence here.
Speed for repeated checks. Jev evaluates independent questions about the same record in parallel. Clarity, relevance and completeness can be assessed in one request. That makes it practical to explore evaluation at the point content is submitted or updated, with response times measured in the actual application.
Low evaluation costs. As of September 2026, TypeSafe lists direct API pricing at $0.042 per million input tokens, with output tokens free. This makes repeated scoring worth considering across a content library. The full workflow budget should also include Gemini proposals, integration maintenance and staff review. See TypeSafe's model pricing.
Verdicts stay within the defined options. Jev returns answers within the categories or scoring levels supplied by the application. That avoids invented verdict labels and unexpected output formats. It can still select the wrong category or assign an unsuitable score, so the team should review accuracy against real examples.
How we use Jev and Gemini at Coworkies
At Coworkies, we now have clear evaluations across space profiles, job ads and blog posts. Those scores help us see which content needs attention according to the criteria we have set for it.
The workflow has two stages: Jev evaluates the content; Gemini creates proposals for improving it. This gives each stage a specific job, with the evaluation providing a focus for the improvement work.
For a space profile, useful criteria might cover how clearly it explains the workspace and who it suits. For a job ad, they might cover responsibilities and candidate requirements. For a blog post, they might cover clarity, relevance and useful detail. These are examples of how a team can define quality for each content type.
Our experience has reinforced a practical point: decide what deserves attention, define how to assess it, then make the result useful in the existing workflow. For operators, the same approach could help prioritize accounts, enquiries and support requests.
Developers can start with TypeSafe's guide to using Jev with coding agents or Cloudflare's Jev model documentation. The surrounding application supplies the records and determines what happens after an evaluation. DataCamp also offers a longer overview of Jev and System One models.
Renewal reviews are a useful starting point
A community manager reviewing 200 accounts has to decide who needs a conversation this week. Renewal dates, recent complaints, attendance changes and notes about growing teams all contribute to that decision.
Consider this scenario for an account:
- Client X: renewal in 18 days.
- Current office: six desks.
- Team size: eight people, according to the latest account note.
- Recent feedback: difficulty finding space when everyone attends.
The capacity mismatch is straightforward to flag with a rule. Jev could help assess the less tidy information around it: a conversation about hybrid attendance, an unresolved complaint, or a request for expansion options.
We could ask it to classify the concern, score review priority against agreed criteria and flag whether a community manager should look at the account. The review screen should show the original notes and dates alongside the result.
A manager could then check attendance plans and discuss a larger office. A writing model could also prepare a proposed follow-up using those records, following the evaluation-and-proposal approach we use at Coworkies.
A review-priority score should have a specific meaning: how soon someone should examine the account. Estimating its probability of leaving is a separate task that requires validation against renewal outcomes.
Sales and support need clear priorities
For sales, useful inputs could include the requested move-in date, team size, available inventory, budget, tour status and recent correspondence. Jev could help classify fit and review priority when an enquiry spans several products or contains ambiguous requirements.
Some follow-up should already be automatic. A completed tour with no proposal after an agreed interval needs a reminder. We would use Jev to assess enquiries that require someone to read the conversation and judge fit or urgency. Every lead should retain a response deadline, including those assigned a low score.
Support triage is similarly concrete. A recurring Wi-Fi complaint, water beneath a dishwasher and a request for whiteboard markers need different owners and response times. Jev could classify the category and urgency, with the application routing each request to the appropriate queue.
Keep established emergency escalation rules in place. Staff should be able to correct a classification, and unclear requests should enter a human review queue.
Treat pricing and location assessments as leads to investigate
Suppose an eight-person office has been vacant for 50 days despite several enquiries. We could ask Jev to classify the objections recorded in sales notes: price, layout, timing, location or insufficient information. The team could compare those results with tour feedback before deciding what to change.
Occupancy and enquiry counts alone cannot establish why a unit is empty. The same caution applies across a portfolio. Low occupancy, high satisfaction and weak lead volume leave several explanations open. A model needs relevant evidence to distinguish them.
Keep invoice balances, occupancy calculations, capacity checks and price scenarios in ordinary software. Our weekly coworking analytics guide gives teams a useful starting set of measures. Consistent definitions also matter across locations, as we explore in the rise of shared data in coworking.
A daily brief needs evidence behind each flag
A useful morning view could surface three items:
- Renewal review: Bright Labs has asked about more space ahead of its renewal.
- Sales follow-up: yesterday's private-office tour still needs a proposal.
- Facilities review: repeated Wi-Fi complaints refer to the same meeting room.
Each item should link to the source record, name an owner and show when the information was updated. Some flags would come from rules; others might use Jev's classifications. The application could present them in one view using fixed templates.
A quiet review queue needs context too. If the support feed failed overnight, the brief should show that gap. An absence of flags is only useful when the team knows what was checked.
Where we would start in a coworking business
Start with one recurring review that takes staff time. Our guide to AI in coworking operations covers how to choose that workflow. Renewal reviews are one option: a sample of 20–50 historical accounts can help a team inspect errors and refine the criteria before a wider trial.
Use only information available before each renewal decision. Later cancellation notes would give the test an answer the team could never have had at the time.
Define three outputs: review priority, concern category and whether human review is needed. Include an insufficient-information category. Compare the results with a simple rules-based queue and an experienced manager's independent assessment.
Then run alongside the current process on new accounts, with staff reviewing every recommendation. Track useful flags, missed concerns, unnecessary alerts and total review time, including corrections. Keep a separate set of cases untouched while refining the questions. A small initial sample helps shape the workflow; reliable performance claims need a broader evaluation.
Use account IDs and only the fields needed for the test. Before sending member records, check the provider's retention and data-use terms. Include integration maintenance and staff review in the cost assessment.
Our view is that Jev is worth considering wherever a coworking team repeatedly evaluates records against clear criteria. At Coworkies, it already helps us identify content that needs attention and gives our improvement workflow a defined starting point. For an operator, the opportunity is to bring that same discipline to the accounts, conversations and requests that deserve action today.
