Blog

  • Learning about the environmental impacts of data centers in Brazil with Rhavena Madeira and André Fernandes

    Learning about the environmental impacts of data centers in Brazil with Rhavena Madeira and André Fernandes

    Data center proposals worth $64 billion have been blocked or delayed by strident civil society organizing across the United States in 2026. But the AI infrastructure boom is showing no signs of slowing down. So if not in the US, where are Big Tech companies taking this vital piece of the AI supply chain? Northeast Brazil has emerged as one lucrative data center construction hotspot. Attractive tax and utility subsidies aside, the region is also a wind and solar energy hub, allowing companies to effectively greenwash their data centers’ socio-environmental impacts. So what does the fight against data center expansion in northeast Brazil look like? 

    As part of their work in our community of practice, Madhuri Karak and Cathy Richards are co-leading a series of events exploring the global socio-environmental impact of data centers. The first edition of the series took place in spring 2026, focusing on the environmental impacts of data centers in the US and the European Union, intending to unpack the development of sustainability metrics within the industry.

    To continue our event series focusing on the environmental impacts of data centers, how they are being regulated and what possibilities exist for resistance, we are inviting Rhavena Madeira (of Filha do Sol, in Piauí, Brazil) and André Fernandes (of IP.rec, in Pernambuco, Brazil) for a virtual roundtable taking place on July 29, 2026, at 12pm ET/1pm GMT-3/6pm CEST. 

    In Brazil, the federal government has been investing in data center expansion, all the while offering little evidence of what the country and its citizens stand to gain from this effort to lure investors. Though the antidemocratic practices of Big Tech companies pushing for data center expansion and the environmental impacts of these large, resource-intensive infrastructures have been pointed out by several civil society organizations, Brazil already has more than 160 data centers and is expected to receive billions in new investments in the sector. 

    We are honoured to be joined by Rhavena Madeira, director at Filha do Sol; and André Fernandes, director at IP.rec, for a conversation about their work against data center expansion in northeast Brazil. Filha do Sol is an organization working in Piauí to uplift feminist and collaborative leadership on the frontlines to regenerate tropical nature and restore climate justice; while IP.rec is an independent center engaged with a focus on the social, ethical, and legal impacts of technological development. IP.rec has recently published this report about AI and the environmental impacts of data centers, as well as this paper on the lack of clarity and opacity that often surrounds conversations about the environment and artificial intelligence.

    About the speakers:

    Rhavena Madeira (she/her) is Co-founder and Strategy Director of Filha do Sol. She is a Brazilian lawyer and project manager with a master’s degree in Human Rights (University of Vienna). For the past ten years, Rhavena has been working to defend socio-environmental rights and a stable climate with the government and civil society. In addition to Filha do Sol, she is also the director of the Center for Climate Crime Analysis Brazil. Rhavena has accumulated experience in managing complex projects and building alliances to defend human rights and climate justice, traditional communities, and local and regional organizations that live and defend their communities in threatened ecosystems, especially in the Amazon and the Brazilian Northeast.

    André Fernandes (he/him) is Founder and Director of IP.rec. André is an Attorney and Professor in Graduate Programs in Law and Technology at UFPE, UPE, and CESAR School. He holds a PhD in Law, with a focus on Artificial Intelligence, the History of Legal Concepts, and Innovation, and a Master of Laws from the Federal University of Pernambuco, with a focus on Legal Theory. He specialized in Artificial Intelligence at the Amsterdam Law and Technology Institute (ALTI) & Vrije University Amsterdam, in AI Policy from the Center for AI and Digital Policy (CAIDP), and is certified in Internet Governance by the Brazilian Internet Steering Committee (CGI.br), DiGI (UCU) and SSIG.

  • Meta just dropped a bomb on chatbot builders. Here’s how it impacts the development and humanitarian sectors

    Meta just dropped a bomb on chatbot builders. Here’s how it impacts the development and humanitarian sectors

    Starting October 1 2026, anyone using WhatsApp as a platform to reach their target audience will be charged by Meta for outbound messages. Until that date, it will have been free to do so within a 24-hour window opened by the user’s last message, meaning that as long as your chatbot (or human agent) was engaging and helpful enough, users seeking support for health, livelihoods, education and psycho-social support could chat for as long as they wanted to (and could afford to themselves). All implementers had to pay for was the cost of the WhatsApp Business Service Provider (BSP) platform, such as Turn.io, as well as any re-engagement messages to wake up ‘dormant’ users. In the near future, any message sent in response to users will be charged at the same rate as ‘utility templates’, whose costs vary region by region.This change has huge scale, sustainability, user experience and impact implications for many social impact organisations using, or thinking about using, WhatsApp to reach people across the Global South.

    How WhatsApp became the channel of choice

    WhatsApp became the digital intervention channel of choice starting in 2018, when its API was released, allowing many of us to seize the opportunity to reach our target audience on the channels they actually use. I was one of very few people designing chatbots back then, experimenting with chatbot copy-writing, design conventions, and figuring out the best way to leverage a fairly restrictive format into a rich experience. Sure, the ‘chatbots’ we created were pretty basic, relying mostly on text, pre-determined decision-trees, and of course, emojis. But they allowed even small NGOs to reach high volumes of users with potentially life-changing information, advice, and support, on an app that was often 0-rated by MNOs, much less prone to connectivity issues than mobile web, and so embedded in users’ lives that it was less likely to be deleted from their phones to preserve memory or battery power. Crucially, users trusted them, loving the human-not-human chat-based format.

    The pandemic turbo-charged the use of WhatsApp by development and humanitarian practitioners, including multilaterals like WHO who deployed chatbots on WhatsApp and other messaging apps to millions of users in over 19 languages, with the help of Turn.io and organisations like Reach Digital Health. At a smaller scale, civil society organisations like the International Fact Checking Network, created a chatbot allowing users to submit health misinformation and get it fact checked in minutes. Today, WhatsApp usage represents 69% of all internet usage worldwide, and 83% of monthly users open it every day. WhatsApp chatbots now feel like old news in the development space ( although we’re still trying to figure out how to measure their impact.)

    GenAI allows chatbots to talk more; Meta pricing is forcing them to talk less

    Whilst launching a chatbot for social impact no longer feels like a novelty, WhatsApp chatbot fever is in no-way diminished. GenAI means that we can finally create *actually* smart chatbots, either run entirely using a Large Language Model, or which include genA-powered components to handle specific conversations requiring more personalisation and flexibility. 

    What this means is that at the exact moment that we are being put under pressure to use genAI to make chatbot experiences richer, more personalised, more flexible, and more accurate, Meta’s new pricing model is forcing us to do the opposite. Turn’s advice to customers using their platform is to ‘reduce the amount of back-and-forth and get users to what they’re looking for as quickly as possible.’ They position this as a user experience improvement as well as a cost-reduction method, and in a general sense, they’re not wrong – many chatbots could do better at being succinct and reducing their feature bloat. But it’s also at odds with the parallel push towards AI-powered, conversational experiences – whose underlying agents will now have to be prompted to be ‘less chatty’, when a big reason for their integration was exactly this. 

    What this will actually cost implementers

    Those with existing WhatsApp services are scrambling to calculate the implications. Informal estimates circulating among implementers range from $500 extra per month, to $4,500 per month to engage with around 10,000 users a month. As one implementer shared: “this is a major constraint to scaling plans and sustainability projections.” For smaller NGOs, what was a cost-effective way of reaching high-volumes of individuals in their target audience with valuable messaging and support, will become prohibitively expensive.

    Projected monthly cost increaseScaleCountryUse Case
    $456010,000 monthly usersSouth AfricaSexuality & mental health
    $36163 million monthly interactionsSouth Africa Learning & Teaching 
    $224518,000 monthly usersArgentinaMental health

    Examples of projected monthly cost increases shared by chatbot implementers

    On the other hand, because exact costs vary per country, some implementers are less concerned. Similarly, for chatbots where the service offering is less dense, and less conversational, the impact may be less significant. As one implementer shared: “I am personally unconcerned. This is because in India, our charges are going from 2 US cents to 7 US cents. Our per-participant cost overall will remain nothing for us.” (their emphases). Additionally, where users come from in the first place matters. Anyone who relies on Meta ads which click-to-WhatsApp for the majority of their marketing will still be able to message users for free within a 72 hour window of the first conversation started. Given that average user lifetime is often only a day, those relying on Meta ads may not be affected at all: basically, don’t panic until you’ve done the maths for your particular set-up.

    The human cost

    Perhaps more seriously, this pricing change may also impact on the role of humans in increasingly automated services. WhatsApp is also used to support human-staffed Helpdesks, with trained operators able to jump in and out of an automated conversation to provide human-led support, including last-mile service provision referrals. These agents are central to safeguarding processes, with users expressing suicidal ideation, survivors of Gender Based Violence, or women going through miscarriages referred as soon as possible to human support. Let’s imagine for a second how it would feel to tentatively disclose spousal abuse, or the loss of a baby, when our interlocutor wants to wrap up the conversation as quickly as possible for cost reasons. 

    How we can push back

    As we speak, implementers leveraging WhatsApp, and the integration platforms who allow them to do so, are discussing ways to fight back, for example by demanding clarity and seeking exemptions for non-profits from Meta; forming a coalition; engaging with country-level regulators; or convincing donors to supplement budgets to cover the unexpected surplus. Whether or not they prevail remains to be seen – the social sector has sought for years with limited success to get Meta to make changes that would enable us to use WhatsApp more effectively in the contexts in which we work. Our collective bargaining power also ebbs away with every cut to international aid. In the meantime, here are some things you can do, right now.

    What you can do – as a donor

    • Get to grips with the cost implications for any organisation in your portfolio using WhatsApp currently and do what you can do to enable them to absorb this unforeseen cost without compromising on service quality.
    • Communicate up the chain about the severity of this situation and the longer term ramifications. Tap into your relationships with regulators, governments and big tech to see how you can make a difference.
    • Make sure anyone proposing to use WhatsApp as an intervention channel is aware of the attendant costs, and can meet them as their intervention scales.
    • Stop expecting too much from a single WhatsApp intervention, unless you can provide the budget to back it up. More features = more back & forth = more budget.
    • Consider the implications for performance and impact indicators. In an age of endless, free, back and forth, conversation length was a measure of success. Not any more. Evaluators will have to disentangle cost-driven user behaviours from organic ones.

    What you can do – as an implementer

    If you already have a WhatsApp chatbot:

    1. Do the maths and figure out how it will actually impact your budget.
    2. As Turn suggests, this is a good opportunity to revisit your User Experience, without compromising on quality or safety. Get help from people with conversation design and UX expertise (see here, here, here and here for staters).
    3. Figure out if you really need that shiny new AI integration.
    4. Talk with your users. Your budget challenges are, of course, not their problem, but talking openly about the implications on the service you provide them could yield valuable insights and unexpected ways forward.

    If you don’t have a WhatsApp chatbot, but were planning one:

    1. Revisit Lean Design first principles and think about what Here I Am calls Minimum Viable Value. What is the one thing that would provide immediate, tangible value, fast, for your users? Build that. And only that, for now. (P.S: we should all be doing that anyway).
    2. Explore other channels. Anyone remember Freebasics? Tom from mySpace? Nothing lasts for ever. Especially for those working with young people, there is almost certainly a platform out there, like Moya Messenger, which will allow you to reach the same volumes of users without the conversational restrictions. 

    *

    This isn’t the end for WhatsApp chatbots, but it’s yet another lesson about relying on corporate behemoths to run essential services. Something to think about before you make GPT, Llama, Gemini or Claude central to your entire intervention – but you know that already. Given their substantial profitability challenges to date, I would bet that the same thing will happen with genAI before too long. 

  • Building AI chatbots people actually trust

    Building AI chatbots people actually trust

    In a recent event hosted by MERL Tech’s Sandbox Working Group, Alyssa Young and Xian Ho of Dimagi walked through several years of work building, testing, and deploying generative AI chatbots for family planning and sexual reproductive health in Kenya and Senegal.

    The project, Open Chat Studio started back before the arrival of ChatGPT as a way to build simple, menu-based chatbots. Once GenAI was in the picture, the team shifted to a different approach. They realized their team, as well as implementing partners, needed a way to create their own GenAI chatbots, so they built OCS and then open sourced it as a digital public good with support from the Gates Foundation.

    OCS sits in between the user and large language model. It’s agnostic to which LLM is being used, whether it’s Anthropic or OpenAI, other providers or open source models found on Hugging Face. It supports a number of channels, including WhatsApp, Facebook Messenger and Telegram. By sitting in the middle, Dimagi can help orchestrate how the different models are responding to user input.

    The session covered some of the important lessons learned during the development of GenAI chatbots in Kenya and Senegal by Dimagi and two organizations working on Adolescent and Youth Sexual and Reproductive Health. 

    In Kenya, the Dimagi team worked with Shujaaz Inc, a Nairobi-based youth brand whose comics, radio, and social channels reach a majority of adolescent girls and young women in the country each month to build the Tubonge chatbot which features two older characters with an existing combined following of nearly 5 million users, were selected through co-design practices as the voices to share the messaging.  

    In Senegal, they partnered with RAES, a Dakar based NGO behind the edutainment series C’est la vie! which already had around 1 million followers across Facebook, Instagram, YouTube and Tiktok. There the chatbot takes on the personality of Marguerite, a midwife, and Dr. Moulaye, a health specialist.

    Key takeaways to consider in chatbot design:

    1. Focus on safety more than the conversation

    When a user messages one of the Kenya bots, the query doesn’t go straight to a conversational model. First, it gets filtered by a crisis classifier which is built to tell the difference between a joke in slang and a genuine cry for help. If it detects an active crisis, it triggers two events:

    • A silent alert to the research team for follow-up
    • Routes the users to a referral with urgent hotline numbers and care options

    If a message is deemed safe, only then will they be sent to the main router which uses chain-of-thought reasoning using the current message and chat history to decide which of the specialized nodes is best prepared to develop a response. The nodes typically include: an advisor (a neutral “probe-then-solve” conversationalist), a referral node (which pulls verified clinics), a roleplay node, a quiz node, and a Shujaaz brand-logistics node. Underneath both instances is retrieval-augmented generation (RAG) drawing only on validated local source material (national health policies, partner content, curated databases, etc)

    2. Listen to users

    Roleplay and quiz features weren’t on the original spec. They came directly out of co-design sessions with young people. Dimagi found that roleplaying gave users a safe simulation space to rehearse hard conversations such as those with a strict parent, a difficult partner, or a judgemental nurse. 

    Users also made requests for sexual and reproductive health (SRH) and mental health content for boys and men as well as offered feedback on the quality of local languages and ways to interact with the chatbot. 

    3. Make sure there is a human in the loop

    Deployment of this chatbot has not been a case of ‘set it and forget it.’ Transcripts are reviewed daily, with a prioritization mechanism that flags conversations needing closer follow-up and a defined human-escalation plan. Dimagi maintains chatbot managers, and that role transfers to local partners as the tools are handed over. Because SRH is sensitive and the technology is new, live implementation includes a safety disclaimer and consent flow, with an age requirement (18 and up) at the front door. As the team shared, the goal is to support not replace human referrals and relationships.

    4. Building it is easy, proving it’s useful is hard

    Dr. Ho’s closing reflection provided the project’s biggest takeaway. Developing an LLM chatbot is simple, but building a genuinely useful one (and keeping it that way) is hard. To mitigate that, the team developed a four part evaluation framework which they use continuously.

    Measuring performance of GenAI Chatbots

     It includes:

    • Regular user testing which measured acceptability and experience through a validated questionnaire. 
    • Expert evaluations with study-team reviewers rating query-and-response pairs on key metrics including accuracy, acceptability, safety, authenticity, empathy, and language quality.
    • Automated evaluations using another LLM as a judge to acquire those same metrics more quickly and widely
    • Head to head comparisons which compared the custom bots to regular ‘vanilla’ bots in blind tests.

    A resource to further explore: 

    Open Chat Studio, Dimagi’s open-source platform for building, testing, and deploying LLM-based chatbots with built-in guardrails. Learn more at sites.dimagi.com/open-chat-studio or browse the code at github.com/dimagi/open-chat-studio.

  • Resumen del evento: lanzamiento del Grupo de Trabajo sobre IA y MERL en América Latina

    Por María de los Ángeles Lasa y Fernando C. Barbosa*. Read in English here.

    El 19 de mayo lanzamos “IA y MERL en América Latina”, un nuevo grupo impulsado por la Comunidad de Práctica de la Iniciativa MERL Tech. La sesión virtual reunió a profesionales interesados en conocer cómo ya se está utilizando la inteligencia artificial en el monitoreo, la evaluación, la investigación y el aprendizaje (MERL, por sus siglas en inglés) en la región.

    A comienzos de este año consultamos a integrantes de la comunidad y encontramos intereses comunes: herramientas prácticas; ejemplos reales y desafíos propios de la región; y preguntas sobre el uso responsable de la IA, incluidas cuestiones éticas, resguardos de seguridad y gobernanza de datos. A partir de estas inquietudes, nos propusimos crear un espacio permanente para abordar estos temas desde las realidades de América Latina y en los idiomas de nuestra región.

    Para el lanzamiento invitamos a dos profesionales que trabajan con la IA en ámbitos distintos pero complementarios: Ana Henríquez Orrego, de Chile, Directora de Auditorías Académicas de la Universidad de Las Américas e integrante de su Observatorio de Inteligencia Artificial en Educación; y Alejandra Lucero Manzano, de Argentina, docente y consultora en planificación y evaluación de políticas de desarrollo y programas de cooperación internacional. La idea era anclar la conversación en sus prácticas cotidianas y abrir un espacio para futuros intercambios.

    Defina un perfil de cargo, no solo un prompt

    Ana Henríquez Orrego presentó una manera sencilla de dejar de buscar el “prompt perfecto”. Cuando un equipo incorpora a una persona nueva, no se limita a darle una única instrucción. Define un rol: qué necesita saber, qué se espera que haga y entregue, qué estándares de calidad debe cumplir y qué límites no puede cruzar. Ana aplica esa misma lógica a la IA mediante lo que denomina un perfil de cargo.

    Este enfoque convierte la elaboración de prompts en diseño de flujos de trabajo. Un asistente de IA útil necesita contexto, un rol definido, productos esperados, criterios de calidad, límites, fuentes autorizadas y un mecanismo para verificar su trabajo. Tomar estas decisiones también permite determinar si una herramienta ofrece suficiente control para un uso profesional.

    Ana también describió tres funciones que puede desempeñar la IA:

    Asistente: realiza una tarea que quien lo utiliza ya comprende y puede supervisar.

    Tutor: ayuda a una persona o a un equipo a aprender mientras realiza el trabajo.

    Personaje: simula a una persona o una situación, por ejemplo, para probar una guía de entrevista o un instrumento de encuesta.

    En el ámbito del análisis de calidad académica, Ana creó dos tutores para apoyar las auditorías de acreditación: uno en NotebookLM y otro como GPT personalizado. Decanos y directores de carrera los utilizan para comprender qué se evaluará, qué evidencias se esperan y cómo prepararse. Las herramientas no realizan la auditoría de manera autónoma, pero ayudan a las personas a participar más activamente en el proceso que ya existe. Ana también utiliza asistentes diseñados para tareas específicas: un asesor curricular, un revisor de resultados de aprendizaje e incluso un corrector de informes que aplica a sus propios borradores.

    Este enfoque también es pertinente para aplicar más allá de la educación superior. En MERL, la IA puede ayudar a organizar evidencia, revisar una matriz de evaluación, comparar documentos o preparar un primer borrador. Sin embargo, quienes la utilizan deben comprender la tarea lo suficiente como para definir el rol y valorar el resultado. Como enfatizó Ana, la IA debe fortalecer y acompañar el trabajo existente, no reemplazarlo.

    Incorporar la IA en el ciclo de evaluación y comenzar a pequeña escala

    Alejandra Lucero Manzano amplió la mirada hacia el ciclo completo de evaluación. La IA puede apoyar la preparación, el diseño metodológico, la recolección de datos, el análisis, la comunicación y el uso de los hallazgos. Sin embargo, el valor de una herramienta depende de la tarea que se le asigne.

    Comenzó con una pregunta sobre las condiciones institucionales: ¿Cuál es el nivel actual de madurez en IA en su organización? Una consultora independiente y una organización de gran tamaño enfrentan oportunidades y restricciones diferentes. La adopción debe partir de esa realidad.

    Al recorrer cada fase, Alejandra mostró cómo la IA puede apoyar la revisión de antecedentes y de términos de referencia, visualizar una teoría del cambio, desarrollar prototipos de pequeñas aplicaciones, simular a una persona encuestada para probar instrumentos y asistir en la codificación cualitativa, el análisis cuantitativo y la comunicación de hallazgos.

    No obstante, una recomendación central fue evitar automatizar todo el ciclo de una sola vez. Para los equipos que se encuentran en una etapa inicial, construir flujos de trabajo pequeños resulta más realista y facilita la verificación. Un ejemplo consiste en elaborar una encuesta con una IA conversacional, solicitar un archivo compatible con KoboToolbox, cargarlo en la plataforma y luego validar el instrumento y su lógica.

    El mismo principio se aplica al análisis. Quienes evalúan deben especificar el enfoque analítico, conocer los documentos fuente y distinguir entre organizar información e interpretarla. La IA puede no captar expresiones regionales, ironías, sarcasmos y otras señales contextuales. Alejandra recomendó solicitar citas específicas y separar el procesamiento descriptivo de la inferencia, de lo contrario, el sistema puede mezclar ambos. El criterio de quien evalúa sigue siendo la principal salvaguarda.

    Y advirtió: un prompt vago produce una respuesta vaga, y ningún prompt puede sustituir la claridad metodológica ni el conocimiento del contexto.

    El uso responsable comienza con reglas institucionales

    Las preguntas de la audiencia llevaron la conversación hacia la credibilidad, la transparencia y la protección de datos. ¿Debe un informe de evaluación declarar el uso de IA? ¿Cómo pueden los equipos generar confianza en productos elaborados con apoyo de IA?

    Ana sostuvo que las instituciones deben establecer las reglas antes de esperar decisiones individuales coherentes. En las instituciones donde trabaja, el uso de IA debe declararse y las herramientas autorizadas se rigen por políticas institucionales. Los datos personales no deben procesarse mediante cuentas gratuitas de uso personal. Los sistemas que manejan esos datos requieren revisión de seguridad, administración responsable y consentimiento informado.

    Alejandra añadió comentarios prácticos: revisar la configuración de privacidad en lugar de aceptar las opciones predeterminadas, separar las cuentas profesionales de las personales, anonimizar los datos antes de utilizar servicios en la nube, y comprender dónde se almacena la información. El consentimiento, la confidencialidad y la gestión responsable de los datos forman parte de una buena práctica de MERL. La IA vuelve estas obligaciones más urgentes y puede hacer que los incumplimientos sean menos visibles.

    La conversación reforzó un punto más amplio: mantener “human in the loop” no es suficiente si carece de la autoridad, el tiempo o el conocimiento contextual necesarios para cuestionar los resultados. Una supervisión efectiva requiere personas que comprendan la evidencia, puedan examinar el razonamiento del modelo y mantengan la responsabilidad por el juicio final.

    Construir un espacio regional para la práctica y el aprendizaje

    De ambas presentaciones surgieron varios principios comunes: comenzar por el problema y el flujo de trabajo, no por la herramienta; definir con claridad el rol, las fuentes, los criterios y los límites del sistema de IA; ajustar la adopción al nivel de preparación de la organización; comenzar con flujos de trabajo pequeños y verificables; y mantener la interpretación contextual, la rendición de cuentas, la transparencia, la privacidad y la gobernanza de datos en el centro del proceso.

    Estas preocupaciones también tienen una dimensión regional. Las herramientas desarrolladas en otros contextos no necesariamente comprenden automáticamente las instituciones, los idiomas, las prácticas profesionales ni los contextos sociales de América Latina. Crear un espacio para comparar experiencias en la región puede ayudar a distinguir qué puede transferirse y qué requiere adaptación.

    Este primer evento inició ese intercambio. El grupo organizará algunas sesiones públicas cada año, todavía funcionará también como un canal abierto para que sus integrantes propongan temas, compartan recursos y se conecten con colegas.

    Puedes ver la grabación del evento de lanzamiento y sumarte a la Comunidad de Práctica para participar en las próximas actividades de IA y MERL en América Latina.

    Agradecemos a Ana y Alejandra por compartir generosamente sus experiencias, y a todas las personas que participaron y plantearon excelentes preguntas. ¡Nos vemos en la próxima!

    * Se utilizó Codex para resumir la transcripción del evento y ajustar la redacción de este contenido. 

  • Event Recap: AI and MERL in Latin America, launching the Working Group

    By María de los Ángeles Lasa and Fernando C. Barbosa*. Lee en español aquí.

    On May 19, we kicked off “AI and MERL in Latin America”, a new group hosted by the MERL Tech Initiative’s Community of Practice. The online session in Spanish, brought together practitioners interested in how artificial intelligence is already being used in monitoring, evaluation, research, and learning (MERL) across the region.

    Earlier this year, we consulted members of the community, and their needs were consistent: practical tools, real examples and challenges from the region, as well as responsible-use questions on ethics, safeguards, and data governance. With that in mind, our aim was to create a sustained space to discuss these issues grounded in the realities of Latin America (and in the languages of our region).

    For the launch, we invited two practitioners working with AI in different but complementary settings: Ana Henríquez Orrego from Chile, Director of Academic Audits at Universidad de Las Américas and a member of its Observatory on Artificial Intelligence in Education; and Dr. Alejandra Lucero Manzano from Argentina, a professor and consultant in the planning and evaluation of development policies and international cooperation programs. The idea was grounding the conversation in their day-to-day practice and opening the space for future exchanges. 

    Write a job profile, not just a prompt

    Ana Henríquez Orrego introduced an easy-to-grasp way to stop chasing the “perfect prompt”. When a team brings in a new colleague, it does not simply give that person one instruction. Instead, it defines a role with quality standards: “Imagine someone is coming to collaborate with you. What would you ask of them, and what instructions would you give so they understand what to do?” Ana applies the same logic to AI through what she calls a “job profile” (perfil de cargo).

    This framing turns prompting into workflow design. A useful AI assistant needs context, a defined role, expected outputs, quality criteria, limits, approved sources, and a way to check its work. Going through those decisions also reveals whether a tool offers enough control for professional use.

    Ana also outlined three roles that AI can play:

    Assistant: completes a task that the user already understands and can supervise.

    Tutor: helps a person or team learn while doing the work.

    Character: simulates a person or scenario, for example to test an interview guide or survey instrument.

    In academic quality assurance, Ana has created two complementary tutors to support accreditation audits: one in NotebookLM and another as a custom GPT. Deans and program directors use them to understand what will be assessed, what evidence is expected, and how to prepare. The tools do not conduct the audit independently; they help people participate more in a process that already exists. She also relies on purpose-built assistants: a curricular advisor, a learning-outcomes reviewer, even a report corrector she applies to her own drafts.

    It matters well beyond higher education. In MERL, AI may help organize evidence, review an evaluation matrix, compare documents, or produce a first draft. But practitioners still have to understand the task well enough to specify the role and assess the result. As Ana emphasized, AI should strengthen and accompany existing work, not replace it.

    Map AI onto the evaluation cycle, then start small

    Alejandra Lucero Manzano zoomed out to the whole evaluation cycle. AI can support preparation, methodological design, data collection, analysis, communication, and the use of findings. The value of a tool, however, depends on the task it is being asked to perform.

    She started with a readiness question: What is the organization’s current level of AI maturity? An independent consultant and a large organization face different opportunities and constraints. Adoption should begin from that reality rather than from an ideal version.

    Walking through each phase, Alejandra showed how AI can support background reviews and terms of reference, visualize a theory of change, prototype small applications, simulate respondents to test instruments, and assist in qualitative coding, quantitative analysis, and the communication of findings.

    Yet her most useful recommendation was to resist automating the entire cycle at once. For teams at an early stage, building small workflows is more realistic and easier to verify. One example was to draft a survey with a conversational AI, request an output compatible with KoboToolbox, upload it, and then validate the instrument and its branching logic.

    The same principle applies to analysis. Evaluators should specify the analytical approach, know the source documents, and distinguish between organizing information and interpreting it. AI can miss regional expressions, irony, sarcasm, and other contextual signals. Alejandra suggested requesting specific citations and separating descriptive processing from inference, otherwise, a system may blend the two. The evaluator’s judgment remains the guardrail.

    Her warning was clear: a vague prompt produces a vague response, and no prompt can substitute for methodological clarity or contextual knowledge.

    Responsible use begins with institutional rules

    Questions from participants brought the discussion to trust, disclosure, and data protection. Should an evaluation report disclose the use of AI? How can teams build confidence in AI-assisted products?

    Ana argued that institutions need to set the rules before expecting consistent individual decisions. In the institutions where she works, AI use must be declared, and approved tools are governed by institutional policies. Personal data should not be processed through free, personal-use accounts. Systems that handle such data require appropriate security review, administration, and informed consent.

    Alejandra added practical safeguards: review privacy settings rather than accepting defaults, separate professional and personal accounts, anonymizing data before using cloud services, and understanding where information is stored. Consent, confidentiality, and responsible data management have long been part of sound MERL practice. AI makes them more urgent and failures less visible.

    The discussion reinforced a broader point that keeping a human “in the loop” is not enough if that person lacks the authority, time, or contextual knowledge to challenge the output. Meaningful oversight requires people who understand the evidence, can interrogate the model’s reasoning, and remain accountable for the final judgment.

    Building a regional space for practice and learning

    Across both presentations, several shared principles emerged: start with the problem and workflow rather than the tool; clearly define the AI system’s role, sources, criteria, and limits; match adoption to organizational readiness; begin with small, verifiable workflows; and keep contextual interpretation, accountability, transparency, privacy, and data governance at the center of the process.

    These concerns are also regional. Tools developed elsewhere do not automatically understand Latin American institutions, languages, professional practices, or social contexts. Creating space to compare experience across the region can help practitioners distinguish what transfers and what requires adaptation.

    This first event began that exchange. The group will convene some public sessions each year but also serves as an open channel for members to propose topics, share resources, and connect with peers in between.

    Watch the recording of the launch event and join the NLP Community of Practice to take part in future AI and MERL in Latin America activities.

    Our thanks to Ana and Alejandra for a generous first session, and everyone who joined and contributed with excellent questions. ¡Nos vemos en la próxima!

    * Codex was used to summarize the event transcript and adjust the text language. 

  • ICT4D 2026: Is Africa’s Response the blueprint for resisting the AI Inevitability narrative?

    ICT4D 2026: Is Africa’s Response the blueprint for resisting the AI Inevitability narrative?

    All roads led to Nairobi from the 20th to 22nd of May 2026 for the ICT4D conference, where sessions centred on the theme “Delivering impact through digital transformation, together.” The MERL Tech Initiative (MTI) was busy leading some amazing sessions as a consortium partner. We led panels on evaluating GenAI challenge funds, resourcing African language NLP for MERL, joined a session on Voice AI, and facilitated a workshop on Competencies for Made in Africa AI in MERL. 

    The week was packed with rich conversations and impressive technology demonstrations. It was fascinating to see how people are advancing geospatial intelligence and climate early warning systems, while still keeping community farmers at the centre of the work. I had stimulating conversations around how technology can support mental health, with some genuinely thought-provoking ideas emerging from those discussions. The innovations on display during the tech demonstrations were truly standout, each offering a fresh and purposeful approach to solving some of  Africa’s problems. Particularly eye-opening was learning about deeper strategies for safety by design, especially around gendered considerations, and how embedding these principles can actually give consumer-facing technology a meaningful competitive edge.

    In this post, I want to share my reflections of the entire experience and some key learning and insights. 

    Minimum Viable Intelligence: Africa’s Existing Asset 

    A lightning talk by Nasubo Ongoma of Qhala was one of the highlights of the event. She posed what may be the most critical question facing the continent: with Africa accounting for less than 1% of global compute, how can we realistically expect to compete with the Global North? More importantly, where do we even begin to build? 

    She reinforced the persistent gap in African language datasets, noting that much is lost in translation when English continues to serve as the default language in AI models. Her argument is that to build AI for Africa, we must first understand and respect the constraints we are operating within, and design around them. If we do this well, we may actually find a way to create and use AI in Africa in a way that is distinctly and powerfully our own.

    Central to this is the nuance of how we can meaningfully layer in existing knowledge systems and cultural frameworks as we build. For technology to truly work in Africa, it cannot be imposed from the outside; community practices and indigenous knowledge must be treated as valid knowledge systems in the development process. AI development in Africa must come with caution, to avoid repeating past mistakes and reinforcing the existing inequalities that place Africa in a perpetual adopter position. 

    Building for and with “community” requires formative research and responsible digital design 

    “Community” was a common and revolving word at this conference. Yet, I kept finding myself in sessions that questioned, reflected on, and interrogated the extent to which communities are actually involved in the technologies built in their name. Practitioners are recognising a critical gap: designing for the sake of technology, rather than designing with the people it is meant to serve. The innovation boom has accelerated this disconnect, prioritising novelty and scalability over meaningful community engagement. 

    A fundamental question remains largely unanswered: how do we actually measure community engagement in tech? Without clear metrics, “community” risks becoming performative, a buzzword dressed up as a value. This is precisely where formative research becomes a valuable asset. Formative research and human- and community-centred design processes aim to ensure that community voices are not an afterthought but a foundation that shapes not just what is built, but the why, for whom, and to what end.

    Data Colonialism 

    While it was encouraging that the organizers elevated this topic to a main agenda item, the execution left much to be desired. I found it deeply problematic for conversations on decolonizing data in the Global Majority to be led by individuals from Silicon Valley and the US. Chief among the issues worth naming explicitly is the positioning of “my experience working in Africa, with Africans, or with African organizations” as a sufficient basis for authority. 

    This framing is problematic because it flattens and renders invisible the African perspective, reducing a continent of immense diversity to a backdrop for external expertise. Our data has been extracted for decades, and now even the story of that extraction, i.e. its meaning, its harm, its trajectory; is being told on our behalf.

    This is what left a bitter taste at ICT4D for me. It raised an uncomfortable but necessary question: why were leading African scholars and practitioners with lived experience not at the center of this conversation? Local voices and in particular, Black African women, feminist scholars, and human rights organizations have been theorizing, documenting, and challenging these dynamics for years. It is only right, then, that the owners of the story be the ones to tell it.

    Building on top of Big Tech models: The need to decentralize power as Africa builds 

    The work of developing AI in Africa cannot happen without directly addressing the issue of power dynamics. It was apparent from the conversations that big tech models such as Open AI’s GPT, Anthropic’s Claude, and Google DeepMind remain the foundational base upon which Africa-focused models are built. While the development of more localized base models in African languages is underway, it is important to continue asking: what alternative models exist, and how are they being built? 

    Equally critical is the question of data ownership and governance — who is collecting data, who controls it, and is it flowing back to the communities that produced it? Community hesitancy and resistance to data leaving the continent reflects a deeper, legitimate concern about extractive practices that have long characterized Africa’s relationship with outside institutions. 

    From the questions raised in various discussions it was a stark reality that vast amounts of institutional knowledge exist in low-resourced African languages, yet many communities lack the technical capacity and resources to transform that knowledge into structured, local-language datasets. 

    What is needed, therefore, is a deliberate decentralization of power, away from big tech and toward smaller organizations, community-led initiatives, and sovereign language models. Alongside this decentralization must come a re-centralization of collaboration: among African institutions, across the Global South, and with aligned partners globally who are committed to equitable AI development. This requires intentional cooperation, shared infrastructure, and a collective insistence that the terms of AI development on the continent are set and led by Africans themselves.

    The Future of Digital Development: Toward African-Led, Locally Grounded AI

    The threads running through ICT4D 2026 converge on a single, uncomfortable truth: localized, Africa led, Africa owned efforts are the only way to reap AI benefits that develop the continent and benefit Africans. Mobile Network Operators (MNOs) offer a rare entry point to this dynamic. They are advancing localization agendas because their commercial viability depends on it. The reach of MNO infrastructure into underserved communities combined with their investment in local language services positions MNOs as consequential actors in the AI localization story. 

    The lesson here is not that MNOs are the solution, but that localization follows where incentives are aligned, and MNOs are strategic partners and enablers for the pathway of locally grounded AI. The hard work here will be ensuring that this localization serves communities and not merely markets – a distinction that matters enormously.

    What is clear is that no amount of community-centered design, sovereign data governance, or African-language AI will gain sustainable ground without an enabling environment at the policy level. The conversations at ICT4D made clear that the public sector holds significant leverage over the terms on which AI enters African contexts through regulation, procurement, data policy, and infrastructure investment. This makes governments an important partner for scaling digital development that is accountable, equitable, and durable. Yet engaging them is rarely straightforward. High turnover in key positions, limited AI expertise, and stretched institutional capacity mean that knowledge and momentum can be lost quickly. Election cycles also bring shifting priorities, and AI itself is increasingly becoming a political issue. These realities don’t make government engagement optional, they make it more necessary to approach governments with strategy.

    Africans are the future of digital development in Africa

    ICT4D 2026 ultimately surfaced that the tools, the talent, the cognisance and the urgency exist. From the question of who tells the story of data colonialism, to who governs the datasets, to who sets the terms of AI development on the continent, the answer must increasingly and unapologetically be Africans themselves. The future of digital development in Africa will be determined by the strength of the ecosystems, institutions, and communities who are leading these efforts. And that is precisely why this moment feels like a blueprint.

    The AI inevitability narrative would have us believe that the path forward is already written, with Africa’s role is simply to adopt, adapt, and keep up.  What emerged from Nairobi is that the African ecosystem is actively choosing the terms of its own digital future and building blocks of an alternative path, one that refuses extraction. Insistence on indigenous knowledge systems, community-centered design, sovereign data governance, African language models, and African-led policy conversations, are not just good practices, they are acts of resistance. If Africa can hold this line, investing in local ecosystems, demanding accountability from big tech, and keeping communities at the heart of innovation, then yes, Africa’s response is the blueprint.

    AI Use Disclosure: The blog utilized Claude Sonnet 4.6 for copyediting, grammar and syntax. The copyedited content was reviewed thoroughly and further edited by the author. The content remains the author’s original ideas and reflects the author’s thoughts and style of writing.

  • Narratives for the Future: What the Global Majority Is Really Saying About AI Autonomy, Regionalization, and Ethics

    Narratives for the Future: What the Global Majority Is Really Saying About AI Autonomy, Regionalization, and Ethics

    The current conversation around artificial intelligence in the Global Majority has, for too long, been framed as a race; a race centred on the rapidly evolving nature of AI and emerging technologies; a race premised on AI’s inevitability and the urgent need for the Global Majority to catch up. This urgency is largely skewed toward regulation and adoption.

    On 25 May 2026, the MERL Tech Initiative’s Community of Practice hosted a conversation on AI autonomy, regionalization, and ethics, led by Chenai Chair, Director at Masakhane African Languages Hub, Wayan Vota, ICT4D strategist and founder of ICTworks, William Tjhi, Deputy Director, AI Products at AI Singapore and Maria Luciano, Tech Policy Analyst. The discussants unpacked what “AI sovereignty”  means for them and who it truly serves. They debated whether it can center people’s rights and well-being, whether it fuels a harmful global AI race, and whether it genuinely challenges Big Tech’s dominance or quietly reinforces it. They also questioned whether locally built AI can serve as a real alternative to giant tech systems, and what sovereignty could offer countries in the Global Majority as a path toward greater independence in knowledge and technology.

    This blog captures the key threads of that dialogue and explains why they matter.

    Why This Conversation Matters for Global Majority Digital Futures

    A narrative circulates across development, policy, and technology spaces: AI adoption is simply the next logical step, and you either get on board or get left behind. Maria pushed back on this framing directly, noting that when people are told there is no choice to be made, the question of accountability disappears. If the AI race is inevitable and adoption is already decided, then the communities most affected never sat at the table where that decision was made.

    This conversation emphasized how narratives shape technological realities. The speakers drew a sharp line among narrative-level framing, the legal and regulatory frameworks that narratives produce, and the real-world infrastructure that ultimately enforces or undermines those policies. When the dominant narrative is “AI is inevitable,” resulting policies tend to treat inclusion as a footnote, and the infrastructure built to enforce those policies carries the same blind spots. Accountability does not disappear in a single event; it erodes layer by layer, beginning with how we talk about technology in the first place. (Read more from Maria on this.)

    Communities across the Global Majority already experience the downstream effects of AI systems built without their input, trained on data extracted without their consent, and deployed in languages never designed to carry the weight of their realities. As Chenai noted, any technology that affects how people access information or exercise their social and cultural rights must also be able to connect with them on their own terms. The choice of language, the choice of modality, the choice to engage or not, these are not technical preferences but foundational design choices. The central question, therefore, is not whether AI is inevitable, but who determines its terms. That question must be present throughout the Global Majority AI conversation: not only at the implementation stage, but at the levels of narrative and design. 

    The Sovereignty Narrative: Local Models Versus Big Tech

    The framing of sovereignty must be honest about what is actually achievable and where real leverage lies. The speakers moved the conversation on AI sovereignty critically beyond its most marketable framing: competing with hyperscalers on foundation model development is not realistic. However, governments and communities can build with genuine purpose and genuine control in targeted areas.

    Wayan put the numbers plainly. Hyperscalers in the United States are projected to spend approximately $700 billion in capital expenditure in 2026 alone, with projections exceeding $1 trillion over the next three years. By comparison, the European Union has announced a $2 billion AI investment, Canada has committed $2 billion to a sovereign AI strategy, and India has allocated $1.5 billion to its AI mission. Even the most well-resourced governments in the world cannot compete with these hyperscalers at their own game. For governments in the Global Majority, the gap represents an entirely different category of problem.

    What made this discussion particularly sharp was the observation that the loudest voices promoting sovereign AI are often those with the most to gain from governments attempting to build it. When Nvidia, the world’s largest chip manufacturer, tours the world’s capitals telling each government they need their own sovereign AI, and that same manufacturer supplies the chips required to build it, the alignment of interests deserves to be interrogated. 

    An alternative exists in small language models. Small language models optimized for specific community needs, edge computing solutions that operate on devices without expensive cloud infrastructure, and domain-specific models in sectors such as health, agriculture, or legal aid represent areas where hyperscalers hold no particular advantage and, frankly, limited incentive to invest. These may be precisely where autonomous and sovereign AI can take root.

    Realistic sovereignty, then, looks like this: build what you can control, contribute where you can influence, and be honest about where you cannot compete. Governments should direct their energy toward what is within reach, thoughtful procurement policy, domestic data governance, investment in local infrastructure, and support for community-led language initiatives, rather than expensive sovereign compute investments that primarily benefit hardware vendors. The conversation also surfaced what sovereignty discussions too often omit entirely: people. Whether the lens is state security, economic competition, or cultural and linguistic preservation, communities—particularly marginalized ones—are rarely the subjects of these conversations.

    Autonomy Is Not Sovereignty: The Practical Distinction

    One of the most clarifying moments in the discussion was the framing of autonomy as distinct from sovereignty in application. The distinction is not limited to definition; it lies in what you build, how you resource it, and who you design it for. Sovereignty, as the speakers explored, anchors its ambitions at the level of the state: national compute, national models, national data governance. Autonomy addresses a more immediate and survivable question: if access to an external system disappears tomorrow, can you still function?

    William offered a practical framework for thinking about whether governments and institutions should build or buy AI capability. The buy-first-then-build logic was endorsed as a reasonable general approach, map your people’s needs through what already exists, then develop targeted domestic capability for what is not being served. The timing of the shift from buying to building matters enormously and varies by context; for some countries and communities, waiting too long means losing the generation of talent and institutional knowledge needed to build at all. (See the full paper here.)

    The goal of building domestic AI capability is not purity of origin but negotiating power. Chenai described the reality of depending on compute infrastructure provided by large technology companies under time-limited agreements. When that agreement ends, the question becomes: what do you have left? If the answer is nothing, then every product built, every community served, and every promise made was contingent on a boardroom decision made without your presence. Autonomy, in this framework, means having enough capability to survive an interruption, even if you choose not to exercise that capability every day. This has direct implications for how organizations think about community guardrails and AI safety in practice. Well-intentioned guardrails tend to protect those closest to the infrastructure. If AI autonomy is to offer meaningful protection for the most marginalized communities, the design of those guardrails must explicitly account for who sits furthest from the system and what reaching them would actually require, in terms of connectivity, cost, language, and trust.

    A hybrid approach emerged as the clearest practical recommendation: work with multiple AI providers, maintain core capability in areas of genuine community need, and build collaborative infrastructure with others who share your goals. No single country or community organization can do this alone. The future of autonomous, community-serving AI in the Global Majority will be built through deliberate cooperation, across institutions, across the Global South, and with aligned global partners.

    The Outlook: Dreaming Carefully, Building Durably

    Several threads converged as the conversation turned to future outlooks.

    First, there was a consistent return to the value of focusing on what is within reach. A well-built, community-governed model in a specific language and domain can serve its people with a depth and responsiveness that no general-purpose model will ever prioritize. Hyperscalers will not build a model that captures the slang, dialect, or specificity of a language spoken by smaller populations, but the people who speak that language can.

    Second, the call for locally defined, alternative evaluation metrics was emphasized. If Global Majority communities develop assessment techniques for their own AI systems that are systematic, not prohibitively expensive, and rooted in what actually matters locally, they gain a form of influence that does not require competing at the infrastructure level. Communities that can assess what works and what does not, by their own standards, change the terms of the conversation.

    Third, the technical complexity of AI is not neutral. Keeping the language of AI governance inaccessible to non-specialists is itself a political choice, a way of ensuring that those most affected by these systems remain outside the room where decisions are made. Education and literacy, not just digital literacy but narrative literacy, the capacity to recognize and contest the stories being told about technology, are as important as compute or data in determining who benefits from AI’s development.

    Finally, environmental sustainability cannot keep being deferred. The Global Majority’s natural resources and human labor are fueling the AI boom, often at the risk of further marginalization and for the benefit of Big Tech. The infrastructure required to build and maintain large AI systems demands enormous resources, frequently extracted in ways that replicate the very patterns of exploitation that communities in the Global Majority are working to escape. A vision of AI that is sovereign and community-serving but environmentally catastrophic is not one worth building toward.

    Closing Reflection

    Collaboration among institutions in the Global Majority, as we continue to interrogate autonomy and sovereignty, is a necessary first step: collectively naming what is at stake and pushing back on narratives that would reduce communities to perpetual adopters of Big Tech’s outputs. Strategizing collectively and charting alternative paths to reclaiming agency in AI is a practical starting point. It may one day produce shared language, shared values, and shared pushback strategies powerful enough to have communities shape how technology is built, countering the mainstream narrative that digital futures depend on Big Tech.

    AI Use Disclosure: The blog utilized Claude Sonnet 4.6 for copyediting, grammar and syntax. The copyedited content was reviewed thoroughly and further edited by the author. The content remains the author’s original ideas and reflects the author’s thoughts and style of writing. 

  • GenAI for Development Needs Its Own Evaluation Standards Before It’s Too Late

    GenAI for Development Needs Its Own Evaluation Standards Before It’s Too Late

    There is a pattern in the history of global development that we keep repeating, and we are repeating it right now with GenAI.

    The pattern goes like this. A new paradigm arrives, a compelling idea about how to address poverty, disease, inequality. Money flows. Governments and foundations and NGOs rush to fund, implement, and report. Programmes scale, and implementation accelerates. But the systems for understanding whether any of it works: how, for whom, and at what cost, arrive late, if at all.

    We have lived through this cycle with structural adjustment, with microcredit, with mobile money, with ICT4D. We are living through it again with GenAI for development. And the stakes, this time, are higher than they have been before.

    How evaluation evolved in development

    The story of how global development came to rely on its current evaluation standards is not one of foresight, but of accumulated failure.

    Figure 1. The evolution of evaluation in global development

    In the decades following World War II, reconstruction aid flowed primarily through bilateral relationships with minimal harmonisation requirements. Development was largely treated as a technical problem, the transfer of capital, expertise, and infrastructure from richer to poorer countries, and the question of whether it was working was answered by outputs: roads built, vaccines distributed, schools constructed. Whether those outputs translated into improved lives was rarely asked systematically, and almost never answered independently.

    The participatory development movements of the 1970s challenged this logic. Theorists and practitioners from the Global Majority, Robert Chambers, Paulo Freire, and many others, argued that development done to communities rather than with them was not only ineffective but potentially harmful. Community agency, local knowledge, and contextual validity were not soft concerns; they were determinants of whether interventions worked or failed. But the evaluation infrastructure of the time had no real way to capture these dimensions. Logframes could count outputs. They could not assess whether the people being served were better off in their own terms.

    The 1980s brought a harder lesson. Structural adjustment policies, imposed largely through conditionality attached to World Bank and IMF loans, shifted responsibility for social services to local governments without providing the resources or capacity to deliver them. The outcomes were, in many cases, catastrophic: collapsed health systems, rising inequality, damaged public institutions. And yet the evaluation systems in place largely registered these programmes as successful against their own narrow indicators. The accountability gap was not incidental. It was structural. The frameworks being used to measure success had been designed by the institutions implementing the interventions.

    By the 1990s and 2000s, the field could no longer ignore the gap between investment and evidence. A series of international commitments: the Rio Earth Summit’s Local Agenda 21, the Paris Declaration on Aid Effectiveness, the Accra Agenda for Action, and the Grand Bargain, tried to address it, establishing principles of ownership, alignment, harmonisation, and mutual accountability. And in parallel, the OECD Development Assistance Committee developed what became the field’s first widely adopted evaluative criteria: relevance, effectiveness, efficiency, impact, sustainability, and later joined by coherence.

    These criteria were not perfect. They were designed for bilateral aid programmes with defined beneficiary populations and stable theories of change. They embed, as critics have rightly noted, a donor-centric logic in which relevance is partly defined by funder priorities, and sustainability is framed around what happens when donor funding ends, implicitly accepting dependency as the baseline condition. The African Evaluation Association, UNEG, and a growing body of scholars from the Global Majority have all pushed back on the limitations of OECD-DAC criteria for contexts that were never consulted in their design.

    But imperfect as they are, these criteria did something essential: they gave the field a shared language for evaluative judgement. They established a floor, a minimum expectation that interventions would be assessed not just on what they produced but on whether they were the right thing to do, whether they worked, and whether the results lasted. That shared language made it possible to compare across programmes, to hold funders and implementers accountable, and to accumulate evidence over time about what works and what doesn’t.

    The AI for development field does not yet have this. And it is spending money as if it does. The lesson from development history is not that evaluation frameworks prevent failure. They do not. The lesson is that with no shared standards for asking hard questions, failure becomes much harder to see, much harder to learn from, and much harder to correct.

    This is the moment GenAI for development finds itself in. The technology is advancing rapidly. Investment is accelerating. But the field has not yet agreed on the standards needed to determine whether these interventions are creating meaningful and sustained value.

    The scale of what is happening

    To understand the stakes, consider the landscape. In the past three years alone, major foundations including Gates, Wellcome, Rockefeller, and Omidyar have committed hundreds of millions of dollars to GenAI for development initiatives. Google.org has launched AI collaborative funds across health, agriculture, and education. USAID, before its decimation in 2025, had made AI central to its digital development strategy. The GSMA has a challenge funds specifically focused on AI for Impact with multiple grantees focusing on different thematic areas. Development finance institutions are beginning to treat GenAI not just as a tool for their grantees but as an investment asset class.

    The tools being funded cut across clinical decision support, agricultural advisory chatbots, legal aid platforms, financial inclusion services, maternal health companions, and educational tutors. They are being deployed in dozens of languages across dozens of countries, serving populations with widely varying levels of digital literacy, connectivity, and trust in automated systems.

    And the evaluation frameworks being used to assess them? In most cases, they are either borrowed from commercial product development, engagement metrics, retention rates, A/B testing, or applied from traditional development evaluation in ways that were never designed for adaptive, technology-mediated interventions. Neither is sufficient. Both can be misleading.

    The AI Evaluation Playbook and its limits

    This gap is beginning to receive attention. The emergence of the Generative AI Evaluation Playbook represents an important step towards building a shared approach for assessing GenAI interventions in development. It is an attempt to answer a question the field urgently needs to confront: how do we move beyond demonstrating that an AI tool functions and begin understanding whether it is actually creating value?

    The Generative AI Evaluation Playbook, developed by Center for Global Development and Agency Fund, is a genuine contribution. It is the most serious attempt yet to build shared evaluation standards for GenAI interventions in the development sector. Its four-level framework (model evaluation, product evaluation, user evaluation, and impact evaluation) does something important: it puts conversations that usually happen in separate silos into a single coherent structure. Normally, tech teams who obsess over model accuracy talk past programme managers who worry about adoption, who in turn talk past evaluators focused on impact. The Playbook forces them to look at each level as a dependency of the others.

    The Minimum Viable Evaluation section is particularly valuable. In a sector where most organisations do not have the resources or expertise for a full four-level evaluation, it provides a practical baseline of “here is the minimum you should be doing,” addressing a real and pressing need.

    The Playbook represents an important foundation. But if the goal is to establish evaluation standards for GenAI in development, three additional dimensions need to be addressed.

    1. Clarifying what kind of evidence each level produces.

    The first issue is conceptual clarity. Before building evaluation standards, the field needs greater precision about what different forms of evidence can tell us, and what claims they can legitimately support.

    The Playbook is often referred to as the “AI Evaluation Playbook,” implying that it covers all kinds of AI, yet it is mainly applicable to GenAI, specifically to chatbots. Other playbooks are likely needed to cover other kinds of (Gen)AI, for example, frontline worker tools, clinical decision-making support bots, computer vision and medical diagnostics, deep tech, and the development of additional language models or wider systems approaches that involve AI.

    In addition, the Playbook describes its framework as four levels of evaluation. But, while it’s common to hear Level 1 referred to as ‘eval’ in the world of AI, not all four levels are considered evaluation in a conventional, development context.

    • Level 1 is model testing and assessment, an engineering activity concerned with whether the AI system produces accurate, safe, and consistent outputs.
    • Level 2 is product monitoring and optimisation, a product management activity focused on whether users are engaging and whether the product performs as designed.
    • Level 3 is user outcome monitoring, which sits closer to programme monitoring than evaluation, tracking whether the product is changing users’ knowledge or behaviour.
    • Level 4 is the only level that constitutes summative evaluation in any meaningful professional sense: an independent judgement, against defined criteria, about whether an intervention is working, for whom, and whether it should continue.

    Calling all four of these “evaluation” obscures what each level can and cannot claim. It risks giving implementing organisations false confidence that running product analytics meets accountability obligations. And it risks giving funders the impression they are receiving evaluation evidence when they are receiving performance data. The Playbook’s own architecture would be stronger, and more honest, if each level were labelled by its primary function: what kind of evidence it produces, and what claims it can legitimately support.

    2. Adding an evaluative framework for judgement.

    The second issue is that measurement and evaluation are not the same thing. The Playbook provides important guidance on what should be measured and assessed at different stages of an AI intervention, but evaluation also requires a framework for making judgements: whether an intervention is relevant, whether it is producing meaningful results, for whom, under what conditions, and whether those results justify continued investment.

    A development evaluator reading the Playbook will find limited guidance on the evaluative criteria or normative framework needed to interpret evidence. Established development evaluation frameworks, including the OECD-DAC criteria, UNEG Norms and Standards, African Evaluation Association Guidelines, do not currently form part of the framework. The result is that the Playbook provides important measurement architecture but does not yet provide the evaluative language needed to make judgements about what the evidence means.

    A related issue concerns sequencing. The framework largely follows product development logic rather than evaluation logic: it begins with assessing model performance before establishing whether the model and the product are the right solution for the right problem. In development evaluation, effectiveness cannot be considered separately from relevance. A chatbot can perform well against technical accuracy benchmarks while still being poorly aligned with what communities need, whether they trust it, or whether it addresses the underlying problem. The field learned this lesson during the ICT4D era: technology can function as intended while adoption remains limited and intended outcomes fail to materialise. GenAI evaluation risks repeating that pattern unless relevance and contextual fit are considered before technical performance.

    3. Ensuring sustainability is treated as an evaluation question.

    The third issue is sustainability. Development experience has repeatedly shown that short-term effectiveness does not necessarily translate into lasting impact. AI interventions are no exception.

    Sustainability is a core development evaluation criterion precisely because experience has shown, repeatedly, that interventions that produce results during the funding period often do not last. For AI tools this is not an abstract concern. The inference costs of running a language model don’t disappear when a grant ends. The institutional capacity to maintain, update, and govern an AI system does not emerge spontaneously. The community trust required for sustained adoption takes years to build and can be destroyed by a single harmful output. An evaluation framework that does not assess whether the conditions for sustained impact are being built is a framework that will systematically miss the most important question about whether AI for development is delivering value.

    What happens without shared standards

    The history of evaluation in development gives a clear answer to what happens when the evidence infrastructure lags behind the investment cycle.

    Without shared standards, the field optimises for what can be measured rather than what matters. Engagement metrics become proxies for impact. Adoption figures become substitutes for evidence of benefit. Funders make portfolio decisions on the basis of outputs rather than outcomes, and the interventions that scale are not necessarily the ones that work; they are the ones that can demonstrate numbers quickly.

    In the absence of evaluative independence, self-reported success becomes the norm. Implementing organisations have powerful incentives to frame their results positively, and without independent evaluation, there is no structural check on this. The history of impact measurement in development is littered with programmes that looked successful in self-reported data and failed in independent assessment.

    Where shared criteria are missing, the field cannot learn across programmes. Evidence from a maternal health chatbot in Kenya cannot meaningfully be compared to evidence from an agricultural advisory platform in India if they are using different metrics, different methods, and different standards for what counts as success. The accumulated investment of hundreds of millions of dollars produces no cumulative knowledge, but a collection of individual case studies that cannot be synthesised into actionable learning.

    This is not a hypothetical future. It is happening now.

    What getting it right would look like

    The GenAI for development field needs evaluation standards for roughly the same reasons that the broader development sector needed the OECD-DAC criteria in the 1990s: money is flowing faster than accountability can follow, and with no shared standards, the field will not be able to distinguish what works from what merely appears to work.

    Getting it right does not require starting from scratch. The Playbook is a foundation worth building on. What it needs is a development evaluation layer: one that adds what it currently lacks but does not duplicate what it does well.

    That layer would include: a new Stage 1 relevance check, grounded in localisation thinking, asking whether the problem has been defined by communities rather than funders, whether GenAI is the right solution, and whether the evaluation framework has been designed with rather than for the people it will affect. It would map each evaluation level to appropriate evaluative criteria, including whether core global and regional evaluation frameworks such as OECD-DAC (adapted), UNEG Norms, AfrEA Guidelines, or a hybrid designed specifically for adaptive technology interventions should apply. It would provide guidance on evaluative independence for accountability claims. It would also include contribution analysis as an alternative causal approach for contexts where RCTs are infeasible, inappropriate, or simply the wrong question. Finally, it would incorporate a sustainability dimension that assesses whether the institutional, financial, and community conditions for sustained impact are being built, besides focusing only on whether the tool is working today.

    (Note: Our team at The MERL Tech Initiative is working on guidance for a new Level 0, which would focus on Formative Research and Digital Design as critical precursors to developing any GenAI ‘solution’ or effort. We’re also examining MERL of GenAI with attention to gender, inclusion, connectivity, community context, and related considerations, as well as evaluation of GenAI integrations and applications that are not chatbots).

    This is a collective task. No single organisation has the disciplinary range to build this alone. It requires technology developers, development evaluators, communities, funders, and researchers working together. It is precisely the kind of collaboration that has historically been difficult to sustain but that this moment requires.

    The history of evaluation in development is, in many ways, a history of learning late. Frameworks and standards emerged after failures had already exposed their absence. The opportunity with GenAI is to do something different: to build the evidence infrastructure alongside the technology, rather than waiting until the consequences of poorly understood interventions become impossible to ignore.

    GenAI may transform development practice. But whether that transformation improves people’s lives will depend not only on what the technology can do, but on whether the field has the discipline to ask where it should be used, for whom, under what conditions, and with what evidence.

    The question is whether we will build those standards before the next wave of scale makes it impossible; or whether, once again, we will learn the lessons only after the costs have already been paid.

    Disclosure: This piece used Claude to assist in drafting a timeline of the evolution of evaluation in global development. The analytical framing, structure, and editorial choices are my own, informed by my professional judgement and experience.  
  • The Made in Africa AI in MERL Framework: The Virtual World Cafe Working Session on July 14

    The Made in Africa AI in MERL Framework: The Virtual World Cafe Working Session on July 14

    We’re excited to take the conversation about Made in Africa AI in MERL from theory to practice and we want you to help shape what comes next.

    On July 14, 2026, at 11am ET / 4pm BST / 5pm CAT / 6pm EAT, the Africa AI Learning Group at the Natural Language Processing Community of Practice (NLP-CoP) will convene a special World Café working session.

    Since publishing the Made in Africa AI in MERL Landscape Study in 2025, we’ve been in ongoing conversation with working group members about what it actually means to use AI in evaluation practice in ways that are grounded, ethical, and locally owned. Most recently, we facilitated a workshop at ICT4D on the competencies required to operationalize Made in Africa AI approaches within evaluation practice. Building on this momentum, we’re ready to take the next step and build a framework together.

    What is this framework, and why does it matter to MERL practitioners in Africa?

    This will be a practical tool and shared reference point that helps practitioners navigate questions like: How do I integrate AI into my MERL work in a way that reflects African values and contexts? What should I be able to do, and what should I expect from the systems and institutions I work with? Currently, those questions don’t have a shared, agreed-upon answer, and this session is where we start working towards building shared agreements.

    What to expect

    There are no formal speakers. No presentations to sit through. Instead, participants will work in small groups with peers from across the continent, debating, co-creating, and stress-testing ideas together. The session centers on values and principles: what should a Made in Africa AI in MERL framework actually stand for, and what competencies, i.e technical, ethical, governance-related, and contextual, do evaluators need to put those commitments into practice?

    By the end of the session, participants will have 

    • Contributed and shaped a  first draft of a Made in Africa AI in MERL Framework
    • Clarity on the competencies that will keep them competitive and grounded as an evaluator
    • New connections with peers who are grappling with the same questions

    We invite you to be part of this important conversation. Anyone who wants to be part of shaping how AI is used in African evaluation practice is welcome to be a part of this. Whether you bring deep technical knowledge, years of evaluation experience, or simply a strong conviction that this needs to get right, we would love to have you participate in this co-creation session. 

    Sign up via the link below to join us for this interesting conversation.

  • We’re hosting a 2-day workshop in Tanzania: AI for MERL Bootcamp 

    We’re hosting a 2-day workshop in Tanzania: AI for MERL Bootcamp 

    This 2-day free in-person workshop will equip participants to think critically about AI, evaluate where it genuinely adds value to MERL workflows, and apply it responsibly. We will cover two key aspects: 1) how to apply AI to support tasks along the MERL lifecycle, and 2) how to conduct MERL of AI-enabled efforts (such as chatbots in health or agriculture).

    Moving beyond theory, participants will leave with hands-on experience using real AI tools, a sharper understanding of the risks, and a clear ethical framework to guide their decisions. 

    Apply here to join us on August 17-18, 2026, in Dar Es Salaam, Tanzania.  Applications close on June 30, 2026.

    Who is this course for?

    This workshop is designed by The MERL Tech Initiative for professionals who work in monitoring, evaluation, research, and learning (MERL) roles across the international development, humanitarian, and social impact sectors. It is particularly relevant for 

    • M&E officers, advisors, and managers seeking to understand AI’s practical relevance to their work
    • MERL leads and directors responsible for guiding organisational data and learning strategies 
    • programme staff who commission or use evaluation findings 
    • researchers and data analysts exploring AI-assisted approaches
    • consultants or independent evaluators wanting to stay ahead of sector trends.

    Level of AI Expertise: Beginner to Intermediate 

    Level of MERL Expertise: Intermediate

    What Participants Will Learn

    By the end of this workshop, participants will be able to:

    • Define key AI concepts and distinguish between different types of AI relevant to MERL
    • Critically evaluate AI tools and assess whether they are appropriate for specific tasks and contexts
    • Identify practical applications of Generative AI across the MERL cycle 
    • Apply prompt engineering techniques to extract reliable, useful outputs from AI systems
    • Recognise the risks of AI adoption, including bias, hallucination, data privacy, and dependency, and apply mitigation strategies
    • Navigate the ethical dimensions of AI use in development settings, including issues of consent, representation, and accountability
    • Make informed decisions about AI adoption (if at all) in their own MERL systems and workflows
    • Adopt frameworks and approaches for monitoring and evaluating the use of AI in programs

    Course Topics 

    • Introduction to key AI terms and concepts 
    • Overview of AI uses and AI use in the MERL sector
    • Introduction to monitoring and evaluation of AI systems
    • GenAI tools and applications for MERL
    • Risks, ethics, and best practices
    • Hands-on experimentation with AI tools for social impact
    • Prompt engineering for MERL practitioners
    • Frameworks for MERL of AI Programmes

    Space is limited, so if you’re planning to be in Dar Es Salaam, Tanzania then, and you’d like to join, be sure to submit your application for the workshop by June 30, 2026.