AI-Powered Software Estimation Shotguns-style ensemble methods is still the best option
While Artificial Intelligence has transformed software development, for software estimation it serves as a powerful enhancement rather than a replacement for established, rigorous methods.
Abstract
While Artificial Intelligence has transformed software development, this analysis highlights that for software estimation, it serves as a powerful enhancement rather than a replacement for established, rigorous methods. The core principle remains that ensemble effort estimation, using multiple independent techniques like a “shotgun” spread, consistently outperforms any single method; and AI/Machine Learning models now contribute a highly valuable, well-calibrated data point to this ensemble. AI’s primary gains are in accelerating scope decomposition, facilitating historical analogy identification from vast datasets, and generating multi-perspective risk assessments, all of which lead to better-calibrated probabilistic uncertainty ranges, ultimately aiding practitioners in delivering the honest, defensible estimates required by complex enterprise programs.
Table of Contents
- The original motivation: hating overtime
- What I did about it, and what I learned
- The Classic Methods are Still Relevant Today
- What the latest AI research has found, and what it confirms
- The old constraints haven’t gone anywhere
- You can’t estimate what you don’t understand
- Finding historical data for comparison, no longer a nightmare
- The expert problem and AI agents as a multi-perspective generator
- Modelling uncertainty properly; we’ve always been too optimistic
- What AI doesn’t replace, and where to stay honest
- So where does AI actually help?
- Conclusion
- References
1. The original motivation: hating overtime
“Somewhere, right now, there is a software project failing” — Roger Pressman [12]
When I started my career at Motorola right after graduating as an Electronic Engineer, two things happened that marked everything that followed. The first was that from that point on, I would live in the world of software engineering, not hardware. The second « and believe me when I say they are deeply connected », was that I became a first-hand witness to the brutal, relentless parade of late nights, weekend shifts, and entire nights spent in the office trying to make an established delivery date. Sometimes we made it. Always at the cost of burnout and overruns that nobody wanted to talk about openly, and contributed to an unhealthy culture of working late.
The root cause, I came to understand, has something fundamental to do with the nature of software itself. Unlike hardware, software, once it exists, can be licensed or copied at near-zero cost « you may still have to pay for the intellectual property, usage rights etc, but it is essentially free to replicate, no need to code it again, recreating what is already existing ». It’s a set of zeros and ones living in a replicable electric reality. Hardware, like everything physical in our lives, can have its design copied, but someone still has to build it. You design a car, a house, a motherboard, you still need a production line. That production model encourages being almost sure with the initial design, i.e. it is free of errors before moving to a production stage, then every error drives a costly learning, more iterations, and economies of scale.
Software projects, by contrast, are almost always about solving problems that have never been solved exactly this way before. Also, we have the luxury of making mistakes early on with almost no cost and starting with no more than an idea. As Fred Brooks said once “Software entities are more complex for their size than for perhaps any other human construct” [24]. This complexity is essential, not accidental, because it reflects the complexity of human institutions, not the simplicity of nature as physics laws do. Once a problem is solved in software, you automate it, commoditise it, sometimes crystallise it into hardware. But the act of building the software? You’re always cutting new ground.
That permanently-behind feeling, the nagging sense that regardless of what kind of project you’re on, there’s never enough time. Roger Pressman, one of the pioneers of software engineering, was equally blunt: “Cost overruns and schedule slippages are a universal phenomenon. They occur in every country and every nationality. No company is immune to this” [12]. That is what sent me down the rabbit hole of software estimation.
2. What I did about it, and what I learned
I decided to go all in. I wrote my master’s thesis on how to improve SW Projects Estimations processes [4] and, simultaneously, focused my statistical Green Belt project in the company I was working back then (i.e. Motorola), on improving the company’s estimation processes from the inside. Two parallel tracks, one obsession: why are we always wrong? and; Can we do better?
I read almost everything that existed on the topic at the time. And the book that shaped my thinking most profoundly was Steve McConnell’s Software Estimation: Demystifying the Black Art [5]. Steve became, in many ways, my idol in this space, someone who managed to bring rigour and intellectual honesty to a subject the industry kept treating as dark magic practised by shamans. « What makes McConnell’s work special to me goes beyond the text itself: I had the privilege of meeting Steve personally on two occasions and attending his estimation course in person. Those sessions were among the most intellectually stimulating experiences of my career ». His core insight, which I absorbed deeply: predicting the future is hard, and the best you can do is generate approximations good enough to contain reality within probable ranges.
The other cornerstone of my learning was Richard Stutzke’s Estimating Software-Intensive Systems: Projects, Products, and Processes [6]; a book I regard, alongside McConnell’s, as the estimation bible for any serious practitioner. Stutzke has a rare ability to make the probabilistic and the empirical feel practically accessible, not just theoretically elegant. If McConnell demystified the black art, Stutzke gave you the engineering rigour to practise it properly.
These two books, combined with essential complementary works, Capers Jones’ Applied Software Measurements [7] on productivity and quality metrics, Putnam and Myers’ Five Core Metrics [8] on the quantitative laws governing software behaviour, the ISBSG’s Practical Project Estimation toolkit [9], Pfleeger, Wu and Rosalind’s Software Cost Estimation and Sizing Methods [10], and Laird and Brennan’s Software Measurement and Estimation [11], formed the intellectual scaffolding of my thesis and internal project research.
McConnell made me think about estimation like this: are you trying to shoot a partridge with a rifle or with a shotgun? The rifle is precise, but if your aim is even slightly off, you miss entirely. The shotgun sends a spread of pellets, you’re modelling uncertainty by design, and your probability of hitting the target goes up considerably. Decades later, the research literature would confirm this intuition empirically: a 2022 systematic performance evaluation by Mahmood and colleagues found that Ensemble Effort Estimation (EEE) techniques generally yield more promising estimation accuracy compared to solo techniques, with machine learning methods being the most frequently implemented in constructing such ensembles [14]. The shotgun was the right weapon all along.
By the time I finished my thesis in 2010, I felt I was at the top of my knowledge of software estimation. And then, as happens, fifteen years passed.
3. The Classic Methods are Still Relevant Today
“[…] It’s still good, it’s still good!” — Homer chasing his BBQ pig [S7, E5, 18th minute]
More than fifteen years later, with AI having changed practically everything else in software, I found myself asking: has something finally been invented that makes all of this obsolete? Has some neural network model, trained on billions of lines of code and thousands of project histories, finally cracked the estimation problem?
The answer, after going back in, reading the new literature, and applying updated techniques across real enterprise programs across many industries, problems and technology stacks at Parser Digital, is no. Not yet. Not in the way that would make the discipline irrelevant.
What AI has contributed is something more like some additional pellets in the shotgun. There are new parametric models based on neural networks and ML approaches that add one more data point, one more calibrated view of where the answer might land. But that is not all, there are many new contributions to the discipline and to the process itself enabled by the new Reasoning Models. They are genuinely useful. But it doesn’t make the other pellets disappear. The fundamentals from McConnell and Stutzke are still the foundation. The classic estimation methods, analogy-based, parametric, bottom-up, expert judgment, Wideband Delphi, are all still relevant, still necessary, and still not interchangeable. What has changed is the toolbox available to apply them, and that’s where AI starts to earn its keep.
The words of Aristotle, which I quoted in my thesis [4] as the opening epigraph for the entire discussion of estimation accuracy, remain the most honest framing I know: “It is the mark of an educated mind to rest satisfied with the degree of precision which the nature of the subject admits, and not to seek exactness when only an approximation of the truth is possible” (Aristotle, 330 BC). Fifteen years of AI progress has not changed that fundamental truth.
4. What the latest AI research has found, and what it confirms
Going back into the academic literature with fresh eyes is a humbling exercise. The research community has been working systematically on AI-based estimation for far longer than the current LLM hype cycle might suggest « and I recall that during our initial estimation improvement project, we launched a parallel effort to leverage the neural networks of the pre-2010 era; however, the results were far from encouraging, likely hampered by the absence of the massive, company-wide historical datasets required for proper calibration », and the picture that emerges is both encouraging and sobering in exactly the right measures.
The first thing worth acknowledging is that the problem was already well-characterised before machine learning entered the picture. The foundational systematic review by Jørgensen and Shepperd, published in IEEE Transactions on Software Engineering in 2007 [23], established after surveying hundreds of cost estimation studies what practitioners already suspected: no single estimation method dominates across all project types, accuracy is highly context-dependent, and combination approaches consistently outperform any individual technique. A follow-up systematic literature review by Wen and colleagues, published in Information and Software Technology in 2012 [13], arrived at the same conclusion specifically for machine learning models: ML methods were competitive with statistical and expert-based approaches on most benchmark datasets, but no single algorithm emerged as universally superior.
Mahmood and colleagues reinforced this picture decisively in a 2022 systematic performance evaluation of 35 studies covering both solo and ensemble ML techniques. Their conclusion was unambiguous: ensemble effort estimation outperforms solo approaches, and ML methods are the dominant building block of those ensembles [14].
The first genuinely landmark paper bringing deep learning to bear on software estimation came from Choetkiertikul and colleagues, published in IEEE Transactions on Software Engineering in 2019 [15]. Their contribution was to train LSTM neural networks on the raw textual content of JIRA issue tracking data from open-source software projects, developer-written descriptions, labels, commit messages, comments, to predict story points directly, without manually engineered features. The results were striking: their deep learning model outperformed every existing automated estimation approach on the tested datasets, demonstrating that the unstructured language of software development artefacts carries more estimation signals than anyone had systematically extracted before.
The significance of this finding is easy to underestimate. It means that the way a team writes about its work: the vocabulary used, the complexity signalled through language, the dependencies described in prose, the ambiguities visible in the phrasing of requirements; contains meaningful predictive power for how long that work will take. Large language models, trained on vastly larger corpora of exactly this kind of language, are the natural and powerful extension of this insight.
Around the same time in 2016, Sarro, Petrozziello, and Harman reframed effort estimation itself as a multi-objective optimization problem, with important results [16]. Rather than optimising purely for prediction accuracy, « the conventional approach », their NSGA-II genetic algorithm simultaneously balanced accuracy, interpretability, and calibration. This resonates directly with what I observe in practice: the most useful estimation output is not the most accurate point prediction but the most honestly calibrated probability range. Optimising for a single number is precisely the trap that produces the overconfident estimates McConnell [5] warned against.
By 2023 and 2024, the research frontier had shifted decisively toward transformer-based models and large language models. The scale of this shift is captured by Hou and colleagues’ systematic literature review of 395 research articles from 2017 to 2024, which identified 85 specific LLM applications across six core software engineering activities, including software management, which encompasses planning and estimation [18]. What was a niche research thread in 2019 had become a major research programme by 2024.
A 2025 systematic mapping study by Rogalski and Smołka specifically focused on LLMs for early-stage software project estimation, cataloguing 30 primary studies and revealing that effort and cost estimation is by far the most frequently targeted application, representing 67 of the broader landscape of related research papers, ahead of schedule estimation (16 papers), project outcome prediction (5), and software size estimation (1) [17]. This is the area where the research community has concentrated its energy, and it is not yet a solved problem. The same study provided important insights into the methodological limitations and challenges of LLM-based estimation approaches, confirming that while the results are promising, the methodology is still maturing and requires careful evaluation.
The challenges are real and cannot be dismissed with optimism. Abbas and colleagues, writing in IET in 2025 [19], identified model generalisation, explainability (XAI), privacy, and algorithmic bias as significant barriers to the reliable adoption of AI-based estimation in practice, noting that sustainable AI deployment in software engineering requires interdisciplinary collaboration, ethical oversight, and clear guidelines to balance technological efficiency with accountability. Bhalla and Jodhka, in a 2025 review of LLMs across the full software development lifecycle [20], added hallucinations, difficulties with long-context reasoning, dataset governance issues, and auditability concerns to this list of open challenges.
These are not theoretical concerns. In practice, an LLM asked to estimate a project with unusual architectural constraints: snapshot-based legacy wrapping, rule externalisation from embedded logic; will hallucinate a confident number derived from patterns it has seen, without flagging that the specific complexity signature of this project is underrepresented in its training data. That is a failure mode that the research community is actively working to mitigate, and it is one that practitioners must keep front of mind.
What the research literature confirms, across every era and every methodological wave, is the same conclusion I reached in my Motorola research in 2010:
“The best results come from ensembles of independent methods [23, 14], from honest probabilistic uncertainty quantification [16], and from grounding estimates in real historical project data [13, 15].” « AI has made all three of these things faster and more accessible. It has not made them optional ».
5. The old constraints haven’t gone anywhere
One thing AI boosters sometimes miss: the business context that makes rigorous estimation necessary hasn’t changed at all.
Highly regulated industries, government programs, large projects that start with an RFP, they all still require an upfront estimate of budget, timeline, and effort. Very few organisations can operate with true Scrum-at-the-letter philosophy, maximising value sprint by sprint without a declared destination. Most need a commitment, fixed price or time-and-material with accountability for the range, before a single line of code gets written.
As DeMarco observed decades ago, with words that resonate as strongly today: “When expectations exceed the possibilities of delivery, projects are destined to a premature death. In such cases, the estimates caused it” [1]. And Fred Brooks was characteristically direct: “More projects have gone awry for lack of calendar time than for all other causes combined […], which is attributable, at least three of the five reasons, to problems related to SW estimations” [2].
I have lived both of these truths firsthand, at Motorola in my early career, and then again across the many large-scale international engagements I have been part of in subsequent companies like Intel, McAfee and finally at Parser Digital. The need for rigorous, defensible estimation has been a constant across all of them, from airlines to financial institutions to public-sector programmes.
6. You can’t estimate what you don’t understand
One pattern I have seen repeat itself across virtually every troubled estimation I have been part of or witnessed: the estimate was requested before the scope was understood in sufficient detail to make any estimate meaningful. As I documented in my thesis [4], drawing on Capers Jones: “organisations routinely violate this concept by estimating costs, effort and duration without even knowing how large the software they need to build will be” [7]. We were being asked to predict the cost of building something whose boundaries, dependencies, and complexity hadn’t been mapped. And as Stutzke put it with characteristic precision: “estimating size is critical to producing credible cost estimates. And predicting it accurately is a very challenging problem” [6].
This is where AI is genuinely transformative, and I say this with conviction, not hype.
Modern LLMs can help you break down a project scope with remarkable speed and intelligence. They can identify architectural dependencies, sketch preliminary designs, decompose epics, flag integration risks, and surface non-functional requirements that human teams under time pressure routinely miss. Bhalla and Jodhka confirmed in their 2025 review [20] that LLMs can effectively analyse project data to create resource estimates and risk assessments throughout the planning phase of the development lifecycle, directly addressing one of the most persistent failure modes in the field.
In one of our recent cases involving a large International Travel company « named as our case study from now on », one of many complex programs in which we have applied this methodology, using LLM-assisted decomposition we quickly identified that what looked like a “tooling extension” was actually a platform build requiring 10–12 backend system integrations, approximately 20 high-complexity ETL pipelines, 300–500 operational legality rules to be manually reconstructed from SME knowledge, 60–75 UI screens, and a unified database schema of 60–80 tables. Getting to that level of decomposition clarity in the early analysis phase would have taken weeks without AI assistance. With it, it took days.
Better scope understanding leads directly to better estimates. That causal chain is simple and powerful, and the research on LLMs for requirements analysis, task decomposition, and dependency identification fully supports it [18, 20].

AI adds pellets to the shotgun. It doesn’t invent it. And after fifteen years, the shotgun is still the right weapon for this hunt.
7. Finding historical data for comparison, no longer a nightmare
“Without data, you’re just another person with an opinion” — W. Edwards Deming
Analogy-based estimation has always been one of the most reliable techniques: find a project similar enough to yours, calibrate for the differences, and use it as a reference anchor. The problem is that finding those analogies has always been brutally difficult.
Tools like SLIM Estimate (from Putnam/QSM) have tried to solve this by maintaining industry-wide project databases, calibrated by sector and project type to finally derive a Productivity Index (PI) which does most of the translation from Size to Effort. COCOMO II did the same but most of the DB used was based on industrial and NASA projects. They usually work, within an order of magnitude. But the best historical data has always been the data from your own organisation, and very few companies have the discipline to collect it in a way it is usable for estimating new projects. As I learned through both my thesis and green belt work, and years of subsequent practice, the collective wisdom of the field is captured in that sharp observation from Capers Jones: “Even with the best estimation tools, performing accurate SW estimates is difficult and complicated. But without good historical information, it is almost impossible” [7].
This is where LLMs change the equation significantly. These models are trained on massive corpora that include industry benchmarks, ISBSG databases [9], published case studies, and thousands of project reports. For the first time, you can query patterns from a vast, reasonably well-calibrated knowledge base and get useful analogy candidates without maintaining a proprietary project history. The RAG-based approaches now emerging from the research literature make this even more powerful: rather than relying solely on the model’s parametric memory, you can ground its reasoning in a curated database of your own past programs, retrieved semantically at estimation time, operationalising, at scale, exactly what Putnam and QSM were attempting with their manually maintained project databases in the 1990s [8].
In our case study, we ran nine estimation methods in parallel, including two analogy approaches (top-down and bottom-up), and seven of them converged to within 10% of each other around approximately 1,000 person-months. That convergence is only achievable when your scope decomposition is solid and your reference data is reliable. And indeed, we used multiple last generation LLM Reasoning Models, from the 3 major providers (OpenAI, Anthropic and Google), to understand and break down the scope.
8. The expert problem and AI agents as a multi-perspective generator
Getting expert judgment on estimation has always carried a structural problem: experts are expensive, opinionated, and anchored to their own experience. Wideband Delphi, structured expert convergence, is designed to mitigate anchoring bias through multiple rounds of calibrated discussion. It works well when you have the right experts and the time to run it properly.
But here’s where AI opens a door that was previously locked: you can now run structurally different LLM agents, each prompted to take a different expert perspective, to ensure no angle gets missed. QA implications, security requirements, regional regulatory constraints, performance and scalability considerations, non-functional requirements, industry-specific compliance. Each of these is an area where human estimation teams, under time pressure, routinely leave blind spots. As Hou and colleagues’ comprehensive survey documents, LLMs now serve identified roles across all six core software engineering activities, including software management [18], which means the repertoire of perspectives available to an AI-assisted estimation team is broader than any single human team could realistically assemble.
This approach echoes the multi-objective framing from Sarro et al. [16]: there is no single right answer in estimation, only a set of trade-offs between different objectives, and AI gives us a practical mechanism to populate that trade-off space systematically rather than relying on whichever expert happens to be in the room.
This matters because, as I argued before and still believe today: as Alfred Pietrasanta once said “Whoever expects a quick and easy solution to the multifaceted problem of resource estimation will be disappointed” [22]. There is no magic technique. There never was. What AI gives us is better coverage of the perspectives we would otherwise miss, which is meaningfully different from giving us the answer.
9. Modelling uncertainty properly; we’ve always been too optimistic
“Uncertainty is a daisy whose petals never finish being plucked” — attributed to Vargas Llosa
One of the most memorable empirical findings in McConnell’s book [5] « one I replicated in my research and focused work » is how systematically optimistic software teams are when specifying their uncertainty ranges. We say “probably between 6 and 8 months” when we should say “probably between 5 and 14 months.” The narrow range feels more professional. It almost always turns out to be wrong.
Patrick Henry’s old observation captured the philosophical core of this problem in a phrase: “I have but one lamp by which my feet are guided, and that is the lamp of experience. I know of no way of judging the future but by the past”. If you don’t have good historical data to anchor your uncertainty estimates, you will always drift toward optimism.
This is where Monte Carlo simulation applied to multi-method estimation outputs becomes most powerful.
For our case study, after collecting the outputs of all nine estimation methods, we ran a Monte Carlo simulation to generate a probability-weighted distribution of outcomes across effort, duration, and cost (not shown here for simplicity). The result was a set of confidence-interval bands that tell stakeholders something far more honest than a point estimate.
The recommended feasibility planning range sits between the 70th and 90th percentile. Not “it may not cost the equivalent to 1000 PM.” But: “with proper risk mitigation and early integration validation, we have a 70% probability of delivering in 25 months at a cost equivalent to 800 PMs, and the ceiling, if the main risk drivers don’t resolve early, is approximately 1000 PM over 30 months.”
That is a different conversation with leadership. And, from experience « where I was on the receiving side too », a healthier one. Because, as Jones observed: “The biggest difference between successful and failed software projects can be encapsulated in two words: no surprises” [7]. Monte Carlo doesn’t eliminate surprises. But it forces you to name them in advance.
10. What AI doesn’t replace, and where to stay honest
The AI/ML model-based estimate in the international travel case came in slightly below the central tendency of all the other methods. Not because AI underestimates systematically, but because the model was trained on historical data that doesn’t well-represent this specific class of architectural challenge: snapshot-based legacy wrapping, rule externalisation from non-exportable embedded logic, and multi-source historical reconstruction. These patterns are underrepresented in most project databases, which is exactly what the research literature would predict. As Choetkiertikul and colleagues demonstrated, deep learning models extract signals from the text patterns they have seen; the patterns they haven’t seen well produce weaker predictions [15]. And as the systematic mapping by Rogalski and Smołka confirmed in 2025, methodological limitations and the challenge of underrepresented project types remain among the most significant open problems in LLM-based estimation research [17].
The broader challenges identified by the research community are equally important to keep in mind. Abbas and colleagues highlighted in 2025 that model generalisation, explainability, and algorithmic bias are significant barriers to reliable AI-based estimation adoption, and that sustainable deployment requires careful interdisciplinary oversight rather than uncritical acceptance [19]. Bhalla and Jodhka added hallucinations and long-context reasoning failures to the list of live risks [20], both of which are directly relevant when an LLM is asked to reason about a novel, multi-system architectural context that falls outside its training distribution. An LLM that hallucinates a confident scope decomposition is more dangerous than no LLM at all. Knowing when to trust it and when to challenge it is a skill that takes time to develop, and it is not optional.
This is what you need to understand: an AI estimation model gives you a very well-calibrated view of what your project would cost if it looked like the average of its training data. The moment your project departs from that average, and genuinely novel programs almost always do, you need expert judgment to apply the uplift that the model can’t see.
Used correctly, the AI estimate is your floor. Expert-informed structural decomposition, domain-specific risk assessment, and Monte Carlo uncertainty modelling give you the full picture.
11. So where does AI actually help?
After fifteen years away from the academic side of this and coming back with fresh eyes and real programs to estimate, here is my honest assessment of where AI earns its place in a serious estimation practice:
Speed of scope decomposition. What used to take weeks of analysis now takes days. This is the single biggest practical gain, and it directly addresses the oldest problem in the field, the organisations that routinely estimate without first understanding the size of what they need to build [7]. LLMs can analyse project data to create resource estimates and risk assessments from early project descriptions, a capability now well-documented across the SE research literature [20].
Access to historical analogies. The problem of finding comparable projects, always one of the hardest parts of analogy-based estimation, and the reason Putnam [8], Jones [7], and the ISBSG [9] spent careers building project databases, is meaningfully reduced when you can query a model trained on massive project corpora. RAG-based pipelines take this further by grounding inference in your own project history.
Multi-perspective coverage. AI agents with different expert lenses systematically cover the blind spots that time-pressured human teams leave open, operationalising the multi-objective framing that Sarro and colleagues showed produces better-calibrated results [16], and drawing on the breadth of SE applications now documented for modern LLMs [18].
Calibrated uncertainty ranges. LLM-assisted range generation, combined with Monte Carlo simulation, makes it much more natural to present honest probability distributions rather than falsely precise point estimates, which, as DeMarco warned [1], are often the root cause of premature project death.
One more calibrated voice. As a parametric method, AI/ML-based estimation adds one more pellet to the shotgun. More coverage. Better probability of getting close to the target. And as both Jørgensen and Shepperd [23] and Mahmood and colleagues [14] established empirically: ensembles outperform solo techniques, every time. The AI model is not the marksman, it is one more pellet in the spread.
What it doesn’t do: replace the discipline of running multiple estimation methods, understanding your scope before you estimate it, presenting honest uncertainty ranges to stakeholders, or validating your critical-path assumptions early. “Whoever expects a quick and easy solution to the multifaceted estimation problem will be disappointed” [22]. That has not changed.
12. Conclusion
Software estimation was never a black art. It was always a discipline, one that requires structural decomposition, honest uncertainty modelling, and the intellectual humility to accept that predicting the future, as Aristotle observed more than two thousand years ago, demands only the precision the subject allows.
AI is making practitioners of that discipline better: faster at scope decomposition, richer in historical analogies, more systematic in perspective coverage, and more natural in presenting probabilistic results to stakeholders. The research community, from the early ML benchmarking of Wen et al. [13] and the ensemble evidence of Mahmood et al. [14], through the deep learning breakthrough of Choetkiertikul et al. [15], the multi-objective reframing of Sarro et al. [16], and the emerging LLM-era systematic evidence of Rogalski and Smołka [17] and Hou et al. [18], is converging on the same conclusion: the tools are getting better, the underlying principles are holding, and the open challenges, generalisation, explainability, hallucination, context sensitivity [19, 20], are real and deserve honest attention.
AI adds pellets to the shotgun. It doesn’t invent it. And after fifteen years, the shotgun is still the right weapon for this hunt.
13. References
Books
- [1] DeMarco, T. Controlling Software Projects. Yourdon Press, 1982.
- [2] Brooks, F.P. The Mythical Man-Month: Essays on Software Engineering. Anniversary Edition. Addison-Wesley, 1995.
- [3] Putnam, L., and W. Myers. Measures for Excellence: Reliable Software On Time, Within Budget. Yourdon Press, 1992.
- [4] Miceli, M. Improvements in Software Estimation Processes: The Case of Motorola Argentina Software Center. Master’s Thesis, Universidad Nacional de Córdoba, FCEFyN, 2010.
- [5] McConnell, S. Software Estimation: Demystifying the Black Art. Microsoft Press, 2006.
- [6] Stutzke, R. Estimating Software-Intensive Systems: Projects, Products, and Processes. Addison Wesley Professional, 2005.
- [7] Jones, C. Applied Software Measurements: Global Analysis of Productivity and Quality. 3rd Edition. McGraw-Hill, 2008.
- [8] Putnam, L., and W. Myers. Five Core Metrics: The Intelligence Behind Successful Software Management. Dorset House, 2003.
- [9] ISBSG. Practical Project Estimation: A Toolkit for Estimating Software Development Effort and Duration. 2nd Edition. Edited by P.R. Hill, 2005.
- [10] Pfleeger, S., F. Wu, and L. Rosalind. Software Cost Estimation and Sizing Methods: Issues and Guidelines. RAND Corporation, 2005.
- [11] Laird, L., and C. Brennan. Software Measurement and Estimation: A Practical Approach. IEEE Press & Wiley-Interscience, 2006.
- [12] Pressman, Roger S. Software Engineering: A Practitioner’s Approach. 2009.
Articles
- [13] Wen, J., S. Li, Z. Lin, Y. Hu, and C. Huang. “Systematic literature review of machine learning based software development effort estimation models.” Information and Software Technology, 54(1), 41–59, 2012.
- [14] Mahmood, Y., N. Kama, A. Azmi, et al. “Software effort estimation accuracy prediction of machine learning techniques: A systematic performance evaluation.” Software: Practice and Experience, Wiley Online Library, 2022.
- [15] Choetkiertikul, M., H.K. Dam, T. Tran, T. Pham, A. Ghose, and T. Menzies. “A Deep Learning Model for Estimating Story Points.” IEEE Transactions on Software Engineering, 45(7), 637–656, 2019.
- [16] Sarro, F., A. Petrozziello, and M. Harman. “Multi-objective software effort estimation.” Proceedings of the 38th International Conference on Software Engineering (ICSE), pp. 619–630, 2016.
- [17] Rogalski, Ł. and Smołka, J. “Large Language Models for Early-Stage Software Project Estimation: A Systematic Mapping Study.” Applied Sciences, 2025.
- [18] Hou, X., Y. Zhao, Y. Liu, Z. Yang, K. Wang, L. Li, et al. “Large Language Models for Software Engineering: A Systematic Literature Review.” ACM Transactions on Software Engineering and Methodology (TOSEM), 2024.
- [19] Abbas, T., S.A. Rathore, A. Turki, S. Khan, et al. “Enhancing Software Engineering With AI: Innovations, Challenges, and Future Directions.” IET, Wiley Online Library, 2025.
- [20] Bhalla, J.S., and M.K. Jodhka. “Leveraging Large Language Models in the Software Development Lifecycle: Opportunities and Challenges.” International Journal of Advanced Research, 2025.
- [22] Laird, L. M. “The Limitations of Estimation.” IT Professional, 2006, Volume 8, Issue 6, pp. 40–45.
- [23] Jørgensen, M., and M. Shepperd. “A Systematic Review of Software Development Cost Estimation Studies.” IEEE Transactions on Software Engineering, 33(1), 33–53, 2007.
- [24] Brooks, F.P. “No Silver Bullet: Essence and Accidents of Software Engineering.” Computer, April 1987, pp. 10–19.
Martín Miceli is a software estimation enthusiast who has spent his career at the intersection of engineering rigour and enterprise delivery. He wrote his master’s thesis on estimation process improvement, and has taught software estimation at very prestigious Universities in Argentina (UNC, UTN). As a C-Level executive at Parser Digital, he has been part of an extraordinary 10x growth journey over the last six years, working with and advising internationally recognised companies and brands across Europe, UK, USA and EMEA on their estimation and delivery processes.
© Parser Digital. All rights reserved.


