Menu

Appier Unveils Risk-Aware Framework to Improve Reliability of Agentic AI Systems

Terry KS 5 months ago

Appier has introduced a new research framework designed to improve decision reliability in Agentic AI systems. The approach evaluates how language models respond under varying risk conditions, helping enterprises deploy autonomous AI more safely and effectively.


SINGAPORE, 11 MARCH 2026 – Appier has announced new research aimed at improving the reliability of Agentic AI systems, addressing one of the biggest challenges in deploying autonomous artificial intelligence within enterprise environments. The study introduces a structured framework that evaluates how large language models (LLMs) make decisions under different risk conditions, providing new insights into how AI systems can operate more responsibly in business-critical scenarios.

The research paper, titled Answer, Refuse, or Guess? Investigating Risk-Aware Decision Making in Language Models, presents a systematic methodology for assessing how AI models decide whether to answer a question, refuse a response, or make an uncertain guess. The approach aims to improve decision reliability in high-risk situations where incorrect responses could lead to significant operational or financial consequences.

The study reflects Appier’s broader research strategy to advance frontier topics in Agentic AI and large language models while strengthening the practical deployment of AI-driven marketing technology. As enterprises increasingly transition from AI copilots to autonomous AI agents capable of executing tasks independently, ensuring the reliability and accountability of AI decisions has become a critical priority.

Recent industry data highlights the scale of this challenge. According to a 2025 survey conducted by McKinsey & Company, 62 percent of organizations have already begun experimenting with AI agents. However, inaccurate outputs and hallucinations remain among the most commonly cited risks associated with enterprise AI adoption.

Appier’s research focuses specifically on addressing two major concerns for organizations: AI hallucinations and decision reliability. As an AI-native company offering Agentic AI-as-a-Service solutions, Appier is working to convert advanced research into practical tools and methodologies that enterprises can use to deploy autonomous AI systems with greater confidence.

One of the key contributions of the study is the development of a Risk-Aware Decision-Making framework that transforms AI decision behavior into quantifiable metrics. Traditional evaluation methods typically measure only whether a model’s answer is correct. However, in enterprise environments, the consequences of incorrect responses or inappropriate refusals can vary significantly depending on the context.

The new framework introduces structured risk parameters to simulate realistic decision environments. These parameters include rewards for correct answers, penalties for incorrect responses and defined costs for refusing to answer. By modeling these trade-offs, the framework enables researchers to analyze whether an AI system is making rational decisions that maximize expected outcomes under specific risk conditions.

Under this methodology, AI models must evaluate their own capabilities, confidence levels and potential risks before determining whether to answer, refuse or guess. The quality of the decision is then measured by how effectively the model maximizes expected reward based on the given scenario.

Using this framework, the research team discovered that many existing large language models exhibit strategic imbalances in their decision-making processes. In high-risk scenarios, models often tend to guess answers even when the probability of being wrong carries significant negative consequences. Conversely, in lower-risk settings, some models become overly cautious and refuse to answer too frequently.

This inconsistency limits both the autonomy and safety of AI agents operating in enterprise systems. The findings suggest that the problem is not solely related to knowledge accuracy, but rather to the difficulty models face when integrating multiple cognitive capabilities into a stable and rational decision strategy.

To address this limitation, the study proposes a Skill Decomposition approach that breaks the decision-making process into three structured stages.

The first stage focuses on task execution, where the model generates an initial response by solving the task. The second stage involves confidence estimation, in which the model evaluates how certain it is about the answer. The final stage applies expected-value reasoning, where the model considers the potential outcomes and risks associated with answering, refusing or guessing.

This structured reasoning process enables AI systems to better evaluate the trade-offs involved in each decision. By separating these cognitive steps, models can produce more stable and rational decisions, particularly in high-risk environments where reliability is essential.

According to Chih-Han Yu, improving the reliability of autonomous AI systems is critical for enabling broader enterprise adoption.

He noted that for Agentic AI to operate effectively in important enterprise workflows, organizations must ensure that AI systems are not only intelligent but also capable of making dependable decisions. By transforming risk awareness into a quantifiable methodology, the research aims to strengthen the foundation for trustworthy enterprise AI.

The findings from the study have already been incorporated into several of Appier’s AI-powered platforms, including Ad Cloud, Personalization Cloud and Data Cloud. These integrations are intended to help enterprises adopt autonomous workflows while maintaining stronger control over AI decision reliability.

Looking ahead, Appier plans to continue advancing research in Agentic AI by combining its proprietary data assets, AI engineering expertise and industry knowledge. The company aims to further support enterprises seeking to deploy AI-driven operations that are not only more efficient but also more transparent and trustworthy.

%d