Skip to main content
    Proof
    Impress Blinds — cost per recorded Google Ads conversion down 62.63%SLS Solicitors — cost per recorded Google Ads conversion down 58.48%FixCare Property — cost per enquiry down 56.36%Rubbish Removal WA — cost per enquiry down 53.18%Floral Cakery — cost per enquiry down 49.82%ILLUMINATE Laser Emporium — cost per enquiry down 48.54%Aussie Plumbing — cost per enquiry down 41.96%Sydney Fence Painting — cost per enquiry down 33.68%Alliance Plumbing — cost per enquiry down 29.6%Gridless Build Solutions — cost per enquiry down 29.16%FacilityWorx — cost per enquiry down 23.62%Cornerstone Roofing — cost per enquiry down 20.47%Council finance — budgeting and workforce designCouncil planning — clearer approvals and reportingCultural audiences — media strategy and creative conceptsCommunity participation — surveys, maps and project pagesPublic-sector research — survey design and analysisPublic art — site studies and visual conceptsRegional identity — visitor guides and wayfinding conceptsCouncil systems — integration and migration planningA pressure washing business — service-led search campaignsA pressure washing business — a clear enquiry journeyA carpet cleaner — Google Ads built around cleaning servicesA roofing company — campaigns for repairs and restorationA CCTV installer — campaigns for security enquiriesA fence painter — search campaigns for specific surfacesA maintenance business — live in 8 weeks, 3 stacks gone41 numbered clauses, published in full5.0 across every Google review$120M+ in media under management250+ active engagements across five countries
    Government & Public Sector

    Give an AI assistant a useful, bounded job

    An AI assistant is easier to evaluate when it has a specific job. “Help staff find approved guidance” is a workable starting point. “Transform the organisation” is not a testable brief. Choose a bounded task, define the review process and judge the pilot against examples the team understands.

    • 23 September 2026
    • 3 min read
    Practical AI assistants and evaluation: map the current work, then connect the information, then support the decision.
    An example approach for a service team answering recurring internal questions.

    A practical guide

    Give the assistant a clear boundary

    • The answer is supported by approved material

      Return a response the reviewer can check

      Keep the relevant source available and test whether the answer addresses the actual question.

    • The material is missing or conflicts

      Flag the gap for review

      The workflow needs a useful route for uncertainty rather than a confident-looking guess.

    • An answer is unsuitable

      Let a person correct or reject it

      Record the issue and use representative questions to decide whether the pilot should improve, expand or stop.

    Pick a task with a clear boundary

    Look for a repeatable information task that consumes time and has an understandable outcome. It might involve finding a passage in approved documents, preparing a first draft or comparing two versions of a document. Write down what the assistant may do and where a person must take over.

    An internal guidance assistant, for example, could help a staff member locate relevant information. It should not quietly become the authority that decides policy or eligibility. That distinction belongs in the brief and in the way the experience is presented to users.

    Prepare the material it will use

    Review the source documents before introducing the assistant. Remove superseded versions, identify the current owner and make conflicts visible. A tool cannot resolve a policy disagreement simply because both versions are available to it.

    Choose a manageable source set for the pilot. Include ordinary questions, ambiguous wording and cases where the answer is not present. The team needs to see how the assistant behaves when information is missing, not only when the correct response appears clearly in the first document it reads.

    Create a review set before the demonstration

    Collect representative tasks from the people who will use the assistant. For each, describe what a useful answer should contain, which source supports it and what would make it unacceptable. Include questions that require the assistant to say it cannot determine the answer.

    Review the output for accuracy, completeness and practical usefulness. A fluent answer can still omit a condition that changes its meaning. Keep the assessment understandable enough that subject-matter reviewers can explain why an answer passed or failed, rather than relying on a single unexplained score.

    Make human review part of the experience

    Users should understand whether they are reading a draft, a summary or a pointer to an authoritative source. Give them a practical way to check the information and report a problem. Decide who reviews those reports and how corrections reach the next version.

    For a drafting assistant, approval may sit with the person responsible for the final document. For a document-search tool, the user may need to read the original passage before acting. Different tasks need different review arrangements; one general disclaimer does not design that work.

    Use the pilot to make a decision

    Compare the pilot with the agreed tasks and observe how staff actually use it. Look at errors and review effort as well as convenience. A useful finding might be that a narrower task works well while a broader one needs better source material.

    The deliverables are a use-case brief, a working pilot, an evaluation set and a clear recommendation. Expansion should follow evidence from that review. The team should also be able to decide that the task needs a simpler tool or a better information process instead.

    Explore our Practical AI assistants and evaluation, related services and Build a dashboard people can check and use.

    What to do next

    Talk to SoudCoh about the service question, the people involved and the outputs your team needs.
    Talk to SoudCoh
    Read next

    Filed under the same desk first. The full index is searchable and filters by reader.

    If you would rather we just did it.

    The briefing above is the reasoning. These are the pages that describe what it looks like as a piece of paid work, including what it costs and what gets reported.

    • Practical AI assistants and evaluationChoose a bounded task where an AI assistant can help people work with information. Define the source material, review responsibilities and examples of acceptable answers before a pilot. Evaluate usefulness and errors with the team, then decide whether to improve, expand or stop the service.

    • Explore Practical AI assistants and evaluationExplore the people, scope and practical outputs involved.

    Apply it to your account

    Reading it is the easy half. Thirty minutes with someone who runs accounts and you leave with a written list of what is leaking on yours — yours to keep either way.

    No pitch deck. No upsell. A real conversation and a written list of leaks.