Expert hands needed if AI is to raise the bar
Andy Woods and Katie O´Connor report on ASC’s November 2025 conference “Beyond the hype: mastering GenAI for real-world insight applications”
The ASC November 2025 conference at The Oval, London, brought together researchers, technologists, and clients to address a pressing question: can generative AI genuinely be trusted to provide better insights within the complex, real-world context of commercial research – and if so, how? Under the banner “Beyond the Hype: Mastering GenAI for Real World Insight Applications,” and with support from principal sponsor ResearchWiseAI and associate sponsor Inspirient, the event focused squarely on practical application rather than promises.
Renowned for its friendly, slightly geeky atmosphere, ASC once again presented a programme blending provocation, practical case studies, and candid reflections on what is effective and what is not.
The morning established the tone with a straightforward challenge to one of the industry’s most enduring workhorses: the survey. Andrew Jeavons (Signoi) argued that “surveys are dead – long live surveys”, reminding us that Likert scales and tick-boxes were a workaround for a world where we could not easily analyse natural conversation at scale. Now that LLMs can work directly with language, the numeric gradation of thought underpinning classic questionnaires looks increasingly artificial. He suggested that richer, more text-based questioning – used to calibrate and tune AI models – may ultimately serve clients better than rigid concept tests.
Maria Colarusso and Andrew Dare (STRAT7 Jigsaw) advanced their research by comparing an AI moderator directly with human researchers in WhatsApp-based interviews. Human moderators still performed better in accuracy and relevance, but participants grew more comfortable with the AI over time – and 77% said they would prefer an AI conversation next time. Their conclusion: AI will not replace qualitative researchers in the short term, but researchers will become designers of hybrid systems, responsible for creating conversational flows, guardrails, and quality checks.
Chris Chowen (Royal Holloway, University of London) introduced vibecoding – using existing AI tools in a quick, improvisational way to sketch and test ideas. From a synthetic survey prototype built in half an hour to a school-safe study assistant and a WebAR “holographic business card”, he demonstrated how AI can turn concepts into working demos in hours instead of months. His message: good inputs, clear goals, constraints, and tone remain key, especially so in vibecoding.
Cameron McAnsh and Andrew Le Breuilly (Purify Intelligence) shifted the focus from data collection to data interpretation, arguing that AI analytics platforms are powerful but often misunderstood. They showed that while LLMs excel at querying dashboards, surfacing patterns, and summarising results, they still struggle with precision analysis and contextual judgement. Their solution was not to sideline AI but to structure it: using MCP connectors to orchestrate controlled queries, enforce definitions, and keep humans accountable for the maths. The message was pragmatic – AI can accelerate analysis, but only when researchers design the workflows that keep them in charge.
Caroline Roberts (University of Lausanne) reminded us that good data begins with good questions. She put LLMs to the test to see whether they could apply established Question Appraisal Frameworks – the rigorous rules researchers use to spot ambiguity, bias, assumptions, or cognitive burden in survey items. GPT-4 matched expert coders on roughly 60-70% of classifications, revealing substantial potential for scalable question review. Yet the gaps were instructive: models struggled with subtle assumptions and context-dependent logic, while humans tended to over-identify problems and apply the frameworks inconsistently. Her conclusion: AI can become a valuable assistant in question evaluation, but only alongside researchers who understand both the frameworks and the limitations of automation.
Ramona Daniel (Twenty3) closed the morning with a sharp warning about “AI slop” – outputs produced when automation outruns understanding. She cautioned that if organisations use AI purely to speed up routine tasks while neglecting researcher development, they risk creating teams that work faster but know less. The danger is not just poor-quality insights but a widening skills gap, as AI handles the creative 20% while humans are left with the repetitive 80%. Daniel argued that the real opportunity is the reverse: using AI to remove drudgery so researchers can build deeper expertise, contextual judgment, and methodological craft. Her takeaway was clear: better AI doesn’t replace researchers, it demands stronger ones.
In the afternoon, attention turned to fully developed agentic systems. Steve Phillips (Zappi) described agents as LLMs combined with prompting and training data, and demonstrated how clients are already creating hierarchies of specialist agents – legal, focus-group, workshop facilitator – that reflect organisational structures. A chocolate-bar development exercise demonstrated how interconnected datasets and coordinated agents can accelerate and enrich the ideation process.
Dan-Martin Hellgren (Research Automators) demonstrated how agentic AI can serve as a synthetic respondent, stress-testing surveys before fieldwork by navigating complex question types and even verifying the programmed questionnaire against the original brief. This prompts challenging questions about bot detection but provides a powerful way to identify errors earlier and more cost-effectively.
Guillaume Aimetti (Inspirient) reminded delegates that LLMs are not built for mathematics – they predict plausible answers rather than perform accurate calculations. Comparing the well-known Pentium floating-point bug with modern models that can be “roughly 50–50” on certain maths tasks, he advocated for a hybrid approach combining deterministic expert systems with LLM-based natural language layers. In Inspirient’s workflow, automated statistical processing (crosstabs, regressions, anomaly detection) is integrated with an LLM, making findings searchable and explorable in everyday language. The outcome is faster analysis with maintained rigour – provided researchers stay alert to hallucinations.
Hasdeep Sethi (STRAT7) provided an insight into such architectures with Crowd Tracks, where planning, signals, and narrative agents work together to generate reports within minutes. His insights were notably realistic: agentic systems tend to produce excess content; images still require human judgement; and depending on a single API or data provider poses a strategic risk. He argued that adoption is influenced as much by transparency and trust as by technological ability.
Threaded throughout the day, a recurring message was clear: AI is not a long-term differentiator; people are. Speakers consistently emphasised that future-proof research teams will be valued less for project management and more for their ability to frame questions, audit data and models, design hybrid workflows, and apply human judgement when automation falls short.
From conversational AI that still requires close supervision, to agentic systems that can drift without oversight, and automated quantitative methods that only succeed when the underlying maths are robust, the conclusion was not that GenAI eliminates the need for expertise, but that it raises the bar.
Conference chair John McConnell concluded the day by emphasising this point. Beyond the hype, mastering GenAI for real-world insight applications will rely on an industry willing to experiment boldly – but also to accept responsibility for how these tools are developed, governed, and utilised.Recordings of the sessions can be viewed here on the ASC YouTube channel.