THE VERY GOOD GUYS / RESEARCH

What we tested.
What we learned.

Can AI answer the question, complete the task and pass the work to the next person or tool? These papers examine real projects and saved tests to find out where things worked and where they broke.

Findings from the work.

01 OCT 2026 / 5 PAPERS

These papers look back at saved results, project files and recorded work. Each explains the evidence, findings and limits. They have not been peer reviewed.

OUR METHOD

Make the claim
match the evidence.

You should be able to see why we reached a conclusion. For new tests, we define a good result and what we will compare it with before seeing the answers.

  1. Define the task.

    Choose the business question, the information we can use and what a good result looks like. Name the person who will make the decision.

  2. Record the whole attempt.

    Save what went in, what came out, what failed and what a person had to fix. Record repeated attempts so they do not look like separate successes.

  3. Check more than the answer.

    Can the next tool read it? Are the facts supported? Did the action finish? How much work did a person still do? Check each question separately.

  4. Keep the limits visible.

    Explain what we directly checked, what someone reported and what we think it means. Include failures and the next test that could change our conclusion.

The current papers use the records available from earlier work. Those projects did not necessarily follow every step of this method. Private source materials remain private.

THE QUESTIONS WE FOLLOW

Across the whole business.

01

Business knowledge and operating instructions

Can recorded knowledge become instructions another person can reliably use?

We study the distance between knowing a business and documenting a process someone else can execute. The work includes interviews, operating manuals, role handoffs and explicit records of unresolved decisions. The meaningful test is whether a person can complete the work, including its exceptions.

  • Recorded process interviews
  • Operating manuals
  • Independent handoff tests
02

Revenue intelligence and customer communication

Does the information used to find and serve customers support the decisions that follow?

We examine prospect classification, call intelligence, qualification and customer communication. A useful system needs clear definitions, supported evidence and an observable next step. We distinguish a successful tool call from an accepted answer, and an identified opportunity from a completed customer outcome.

  • Website classification
  • Call opportunity definitions
  • Voice-agent acceptance scenarios
03

Operational agents and verifiable actions

What must be true before an assistant can produce a dependable business result?

We investigate data coverage, business definitions, permissions, delivery and recovery across operational systems. The work connects conversational assistants with deterministic controls. Each capability needs a bounded acceptance test, including what happens when a source is unavailable, a request repeats or an action remains unconfirmed.

  • Warehouse reconciliation
  • Action-policy checks
  • Delivery and retry behavior
04

Evidence quality and business analysis

Is the available information sufficient for the decision being made?

We study incomplete records, uncertain proxies and the assumptions inside business metrics. Methods include source inventories, explicit fallback rules and sensitivity analysis. The goal is to show what a dataset supports, what it leaves unresolved and which additional observation would make the decision stronger.

  • Document coverage inventories
  • Ranking under uncertainty
  • Deterministic metrics with model explanations
05

Creative production and fulfillment

How does a generated asset become an accepted, deliverable product?

We examine consistency across generated video, personalized publishing and production workflows. Attractive previews are one stage. Reference fidelity, revision control, technical checks, approval and fulfillment each need their own evidence. We retain failure cases to understand where generation ends and dependable production begins.

  • Video geometry review
  • Personalized-book rendering
  • Proof and fulfillment boundaries
06

Adoption, evaluation and human decisions

Can people understand, correct and use the system in the work that matters?

We treat adoption as part of system performance. Training, decision authority and the customer experience influence whether a technically functioning tool is useful. Our research questions include who reviews an output, how disagreements are resolved and whether the workflow fits the people expected to use it.

  • Operator training
  • Shared evaluation tasks
  • Human review and customer fit

A QUESTION IN YOUR BUSINESS?

Let’s find out
what is possible.

Bring us a task, a set of numbers or a question. We can run a small test and show you the results, likely costs and a clear next step.

Find your first projectSee how we work together