A controlled experiment with 22 developers, 110 real-world tasks and three LLMs proposed Hybrid Intelligence Effort, combining model work with human oversight. Story Points retained partial value, but validation and corrective intervention dominated effort; HIE dimensions reportedly explained about 72-80% of observed variance. Results may vary by model, team and task.
Key findings
- Story Points retained partial explanatory validity but missed dominant effort sources in LLM-assisted work. HIE dimensions reportedly explained roughly 72-80% of observed effort and reduced systematic error; human validation and corrective intervention outweighed artefact-level characteristics.
Why this matters globally
If replicated, the framework could improve staffing, budgeting and scheduling for AI-assisted projects and reveal when coding-time savings are offset by verification work.
Thai researcher contribution
Arfat Ahmad Khan of Khon Kaen University's College of Computing is the Thai-affiliated co-author in this empirical software-engineering and human-AI study.
Limitations to consider
Twenty-two developers are a small sample. Three models and selected tasks may not represent organisations; prompting skill, tools, domains and quality bars affect effort; models evolve quickly; and interaction measures may proxy task difficulty rather than cause effort. Full measurement definitions require review.