Research
Can an agent turn fragmented intent into a completed group, without breaking a hard rule?
Doubles runs a small, applied coordination research effort alongside the product. It is product research, not academic research: the aim is to know what is safe to automate.
Research (A question we are working on. No result is claimed.) Everything below is a question, a definition or a design. No results, scores or benchmarks have been produced.
01Research question
The question, and the ones underneath it.
Can frontier agents reliably turn fragmented human intent into completed real-world group actions while respecting hard constraints?
- When should model reasoning hand off to deterministic optimization?
- How should hard and soft human preferences be represented?
- How should the system recover when one participant cancels?
- How do we measure coordination quality?
- How much clarification is necessary before an agent can act?
- When should an agent ask a question rather than infer a preference?
- How can a system avoid repeatedly grouping the same people?
- How should uncertain skill levels affect matching?
- How do we measure unnecessary intervention?
02System hypothesis
Propose, then validate.
Research (A question we are working on. No result is claimed.) A system in which a model proposes and a deterministic validator disposes can make hard-constraint violations at execution impossible by construction. If so, the remaining research is about proposal quality: completion, clarification and intervention. Our working belief, to be tested, is that most of the value is there rather than in validation.
This is a hypothesis. It could be wrong, for instance if validation turns out to reject so many proposals that completion suffers.
03What we measure
Nine evaluation dimensions.
| Dimension | The question it answers |
|---|---|
| Hard-constraint violation rate | How often does a proposed or executed action break a rule that cannot be broken? |
| Group completion rate | Of the groups the system sets out to form, how many reach four confirmed players? |
| Human clarification rate | How often must a player be asked a question before the system can act? |
| Cancellation recovery rate | When a confirmed player drops, how often is the group restored without a human? |
| Manual intervention rate | How often does a human operator have to step in? |
| Matching stability | Does a small change in input cause a large, unjustified change in the proposed group? |
| Preference satisfaction | How many stated soft preferences are honoured in the final group? |
| Time to complete coordination | How long from first request to a confirmed game? |
| Tool-call failure recovery | When a tool call fails or times out, does the workflow recover safely? |
Hard-constraint violation rate is the gating metric: no autonomy is extended while it is above zero on the test set. Everything else is traded against it, never the reverse.
04Synthetic test cases
Scenario families.
Research (A question we are working on. No result is claimed.) Scenarios are written by hand and generated variants, each with a known correct outcome, so grading does not depend on opinion.
- Contradictory preferences
- A request whose own conditions cannot all hold.
- Ambiguous time and place
- Daypart words, “around” a city, “the other side”.
- Uneven skill pool
- Pools where the obvious group has a wide skill spread.
- Incomplete groups
- Three compatible players, or five for four seats.
- Cancellation cascades
- A drop, then a failed replacement, then another drop.
- Stale inventory
- A court that changes state between proposal and action.
- Repeated pairing
- Pools that tempt the system to group the same people again.
- Tool failures
- Timeouts, partial success and duplicate calls.
- Hostile text
- Requests that contain instructions aimed at the system.
05Real-world loop
How real coordination feeds the evaluation.
Padel games in Delhi NCR are arranged by hand today. That work is the source of real phrasing and real failure cases. The intended loop, which is Planned (Intended. Not built.), is:
Collect, with consent
Keep structured intents and outcomes from real coordination. Do not keep more free text than the evaluation needs.
Replay
Run those cases through the agent offline and grade against what actually happened.
Shadow
Let the agent propose alongside the human organiser without acting. Compare proposals.
Extend autonomy narrowly
Only for action classes where the hard-violation rate has been shown to be zero on the test set.
06Failure taxonomy
Seven ways coordination fails.
- A. Interpretation
- The model misreads a time, area, skill band or flexibility.
- B. Constraint
- An action would break a hard rule. Must be caught by the validator; counts as a violation if it is not.
- C. Coordination
- A valid group that is a poor one: unstable, repetitive or needlessly unmatched.
- D. Process
- Asks too much, asks too little, or involves a human unnecessarily.
- E. Tooling
- Timeouts, partial writes, duplicate actions, stale reads.
- F. Adversarial and privacy
- Injected instructions, over-retention of text, data shown to the wrong player.
- G. Human factors
- No-shows and late changes that no system can prevent, only absorb.
07Open questions
What we do not know.
- How many requests a human organiser can take before shadow-mode comparison becomes statistically meaningful.
- Whether players prefer being asked one question or being offered two options.
- How much a model’s uncertainty about skill should widen the allowed spread.
- Whether the architecture transfers to activities with different group sizes. Not tested.
The implemented pieces so far are the intent boundary and the validator; see Technology and Evidence.