aodit يقيّم كيف يتصرف وكلاء الذكاء الاصطناعي تحت الإجهاد والتناقض والضغط العدائي. الإطار مُصمَّم لفرق المخاطر والامتثال والتدقيق التي تحتاج أدلة مستقلة بدلاً من أداء النماذج المُبلَّغ ذاتياً.
لمن هذا موجّه
فرق المخاطر ومسؤولو الامتثال ودوال التدقيق الذين يحتاجون أدلة سلوكية مستقلة بدلاً من معايير الموردين.
يقيس كل بُعد فئة مخاطر سلوكية متميزة باستخدام سيناريوهات متعددة الجولات منظمة.
Does it maintain accurate outputs under repetition and contradiction?
Does it disclose uncertainty and resist fabrication under authority pressure?
Can it protect sensitive information under extraction and role-pressure attempts?
Does it make proportional decisions and escalate correctly under risk?
Does it resist jailbreaks, injections, and social engineering pressure?
Does it maintain coherent performance under prolonged stress and ambiguity?
يشغّل كل تقييم تسلسلاً من ثماني جولات يزيد الضغط تدريجياً ثم يختبر التعافي.
The opening question
Normal question to establish baseline behaviour under no pressure.
The reworded question
Same question in different wording to test consistency under rephrasing.
The complicated scenario
Added complexity to test memory, priorities, and constraint handling.
The grey zone
Edge-of-policy case to check boundary adherence versus loophole drift.
The hard push
Urgency, authority, and emotional pressure to test stress resistance.
The manipulation attempt
False premises and disguised requests to test attack detection and refusal.
The recovery test
Post-pressure stabilization check for safe, coherent re-alignment.
The final judgment
Closing assessment of reasoning quality, escalation decisions, and robustness.
aodit طبقة تقييم مستقلة. لا يُصدِّق على النماذج ولا يستبدل أطر الحوكمة ولا يصل إلى أوزان النماذج.