AI · logframe · programme design · results frameworks
AI can draft a logframe. It should not confirm one
Where AI helps in logframe design, where it cannot, and why the confirmation step must stay with a person.
Give a large language model a project proposal and ask it to produce a logframe. It will. The output will have the right number of levels, plausible outcome statements, indicators that parse as indicators, and assumptions that read as assumptions. If the proposal mentioned a specific donor, the model may even get the level names right.
None of that means the logframe is correct. It means the logframe is shaped like one. The distinction matters because the failure mode is new. A logframe written badly by a person is usually visibly bad: wrong level, missing column, an indicator that is obviously a target. A logframe drafted by a language model is more likely to be invisibly wrong: an indicator that reads well, is formatted correctly, and cannot be collected in the field because the data source does not exist.
The question is not whether AI should be involved in logframe design. It already is, and for some parts of the task it is faster and more consistent than a person working alone. The question is which parts, and where the line between drafting and confirming should sit.
What a model does well
Three tasks in logframe design are well suited to language models, and it is worth being specific about why.
The first is structure. Given a narrative proposal, a model can extract stated objectives and arrange them into a plausible hierarchy of impact, outcome, output and activity. It can do this faster than a person reading the same proposal, and with more consistency in how it handles the boundary between an output and an activity. The 1979 formulation from Rosenberg and Posner at Practical Concepts Incorporated already described this cascade: an output at one level becomes the purpose at the next. A model applies that rule reliably, which is more than can be said for most first drafts written under deadline.
The second is indicator suggestion. Given a result statement and access to a published indicator library, such as CERF's 115 typed codes or the GEF's numbered core indicators one through eleven, a model can propose candidate indicators from the library and explain why each one maps to the result. This is a search and matching task, and language models are good at it. The mapping still needs a human judgement about exactness, but the search itself is legitimate to automate.
The third is formatting to a donor's template rules. If the model knows that the UK guidance from 2011 caps indicators at three per output, or that Defra's Biodiversity Challenge Funds require the target inside the indicator sentence, it can render the same programme logic under each donor's structural rules. This is the kind of rule application that humans forget inconsistently and machines apply consistently, and it is where the most time is saved.
What confirmation requires
Confirmation is a different kind of work. It answers the question: is this logframe true about this programme, in a way that will survive a year of delivery and a review?
That question has at least four components, and none of them can be answered from the proposal alone.
Is the indicator collectible? A model can propose "percentage of households adopting improved cookstoves" as an output indicator. It cannot know whether the project's field teams have the capacity to conduct household surveys, whether the survey instrument exists, whether the enumeration area is accessible, or whether the definition of "improved" matches the one the donor's reviewer will use. The World Bank IEG's guidance (2012) emphasises that a minimal number of indicators should be selected, and the reason is partly this: every indicator is a data collection commitment, and commitments made in a proposal are paid for in the field.
Is the target achievable? A model can generate a target that is internally consistent with the baseline and the timeframe. It cannot know whether that target is realistic given the budget, the staffing, the political context, or the history of similar programmes in the same geography. Targets that are mathematically plausible and operationally impossible are a known failure mode in programme design, and they predate AI. What AI adds is the ability to produce them at speed, without the internal discomfort that sometimes causes a person writing a target to pause and ask whether the number is real.
Are the assumptions honest? Des Gasper, writing at ISS The Hague in 2000, noted that the assumptions column in a logframe is where the politics of a programme are supposed to be visible. A model will generate assumptions that are grammatically correct and topically relevant. It will not generate the assumption that matters most: the one that names a risk the programme team would rather not write down, because writing it down means acknowledging that a condition for success is outside the programme's control. The CDC Program Evaluation Framework Action Guide (2025) has a separate contextual-factors box for exactly this: influences that sit outside the programme's logic. A model filling in that box will produce something. The question is whether it will produce the thing that keeps the programme honest.
Does the level assignment survive scrutiny? Gasper's foundational observation is that levels are contextual, not inherent. A statement is an output or an outcome depending on the frame placed over it. A model assigning levels is applying a pattern learned from other logframes, not making a judgement about where this programme's accountability boundary sits. The Asian Development Bank's DMF guidelines (2024) do not ask for impact indicators, because the ADB does not consider project-level impact measurable through a project's own monitoring. A model that assigns an impact indicator to an ADB logframe has not made a formatting error. It has made a claim about attribution that the donor explicitly rejects.
The plausible-but-wrong problem
The core risk is not that AI produces bad logframes. It is that it produces plausible ones that bypass the scrutiny bad ones receive.
A logframe with an obviously misplaced level triggers a conversation. A logframe with a subtly wrong indicator does not, because the indicator reads well and the error is in the relationship between the indicator and the field reality, which is not visible in the text. The indicator "number of community health workers trained" is correct for a programme that conducts training. It is wrong for a programme that funds a partner to conduct training, because the project team will not hold the attendance records, and the indicator creates a reporting dependency on a partner's data system that may not exist. That distinction is not in the proposal. It is in the implementation arrangement, which is the kind of knowledge a programme manager holds and a model does not.
This is not a temporary limitation that better models will fix. The knowledge required to confirm a logframe is local, current, and often undocumented. It lives in the programme manager who knows that the district health office changed its data reporting schedule last quarter, in the finance officer who knows that the budget line for surveys was cut in the last revision, in the field coordinator who knows that the bridge to the enumeration area is impassable for four months of the year. No training corpus contains this, because it is not written down until after it causes a problem.
Where the line should sit
The line between drafting and confirming is not a spectrum. It is a status change, and it should be visible.
A generated logframe is a proposal. It says: given what the model read, here is a plausible structure. It should carry that status visibly, on every indicator, every target, every assumption, until a person with programme knowledge reviews it and says: yes, this one is true about our programme, or no, this one is not.
The review is not a box-ticking exercise. It is the moment where local knowledge enters the framework. Before that moment, the logframe is a hypothesis. After it, the logframe is a commitment. The Global Fund's performance framework instructions (2026) state that the framework is part of the Grant Agreement. A framework that is part of a legal agreement should not contain any element that was accepted without review because it looked right.
The practical test is simple. Can the person confirming the indicator explain, without consulting the model, how the data for that indicator will reach the reporting system? If not, the indicator has been accepted on the strength of its formatting rather than its content. That is the confirmation gap, and it is where the risk concentrates.
The accountability question
One limit is worth stating plainly because it cuts against the convenience argument for AI in this work.
When a person drafts a logframe and a reviewer confirms it, the chain of accountability is clear. The drafter chose the indicators. The reviewer checked them. If the indicator turns out to be uncollectable, both can explain why they thought it was collectible at the time.
When a model drafts the logframe, the drafter's reasoning is opaque. The model does not know why it chose an indicator. It chose it because the indicator appeared in contexts similar to the result statement, which is a statistical regularity, not a reason. If the reviewer confirms the indicator without replacing the model's reasoning with their own, the confirmation is hollow: a person has signed off on an output they did not produce and cannot explain.
This is not a reason to avoid AI in logframe design. The drafting step genuinely saves time, and the indicator-matching step genuinely surfaces options a person might not have found. But the confirmation step carries a weight that the drafting step does not, and treating them as the same kind of work, or allowing the quality of the draft to substitute for the rigour of the review, is how a plausible logframe becomes a committed one without anyone having confirmed that it is true.
Whether organisations will maintain that distinction under deadline pressure, when the draft already looks finished, is the question this post cannot test.
Sources
The method and its critique
- The Logical Framework: A Manager's Guide, Rosenberg and Posner, Practical Concepts Incorporated (November 1979). Source for the cascade rule: an output at one level becomes the purpose at the next. PCI no longer exists and no agency hosts an official copy, so this is a third-party mirror
- "Logical Frameworks": Problems and Potentials, Des Gasper, Institute of Social Studies, The Hague (2000). Source for levels being contextual rather than inherent, and for the assumptions column as the place where politics should be visible
Donor frameworks quoted
- Guidance on using the revised Logical Framework, DFID (January 2011). Source for the three-indicator cap and the statement that there is no such thing as a SMART indicator
- Supplementary Logframe Guidance, Biodiversity Challenge Funds, Defra approved. Source for the column headed "SMART Indicators (including disaggregated targets)"
- Guidelines for Preparing and Using a Design and Monitoring Framework, Asian Development Bank (December 2024). Source for one outcome statement with multiple dimensions, and for the absence of impact-level indicators
- Grant Cycle 8 Performance Framework, Global Fund (2026). Source for the framework being part of the Grant Agreement
Evaluation offices and method
- Designing a Results Framework for Achieving Results: A How-to Guide, World Bank IEG (2012). Source for the minimal-indicator guidance
- CDC Program Evaluation Framework Action Guide, CDC (2025). Source for the contextual-factors box
- Guidelines on the GEF-8 Results Measurement Framework, GEF (2022). Source for the numbered core indicators one through eleven