All research questionsFIDE RESEARCH QUESTION / 05 01 / WHAT THE RECORD CAN SAY 02 / WHAT REMAINS OPEN 03 / A MORE DECISIVE TEST OBSERVATIONS A STUDY WOULD NEED NEXT RESEARCH QUESTIONDoes a convincing report faithfully represent its evidence?
When does an earlier evaluation stop being informative?
A model can remain unchanged while its instructions, inputs or available tools change. Results describe the configuration that was actually evaluated.
Protocol developmentFide is developing a comparison between limited retesting and complete reruns for malware-report analysis.
Can a small, targeted retest reliably detect a meaningful decline? Software checks of a pipeline cannot answer this empirical question.
Collect paired results before and after a specified setup change. Compare limited retest signals with the full rerun, separating changed responses from changed scoring.
PROPOSED TEST · NO RESULT CLAIMEDWhat would change the answer?
- Missed regressions
- False alarms
- Test usage
- Scoring effects
LINKED PUBLICATIONS
Follow the evidence and its limits.
Source type and limitation are kept visible. A research plan is not an observed outcome.
01
RESEARCH PLAN · Fide AI
02When Security Evaluations Go Stale
Protocol development and software checks are not model-performance findings. Live-model experiments remain ahead.
BENCHMARK · Meta & CrowdStrike researchers
03CyberSOCEval
Report-based analysis does not directly measure live incident response or the effects of autonomous actions.
TOOL · UK AI Security Institute
04Inspect: a framework for language model evaluations
An evaluation framework provides infrastructure; the validity of conclusions still depends on task and scoring design.
BENCHMARK · Inspect Evals contributors
CyberSecEval 4 in Inspect Evals
Autonomous-uplift and autopatching prototypes are excluded. Reported subset results do not establish deployment readiness.