Evaluation & testing
Project Moonshot
Apache-2.0 LLM evaluation toolkit from Singapore's AI Verify Foundation that combines benchmark testing across safety and performance metrics with manual and automated red-teaming, offers guided workflows for IMDA's starter kit for LLM app testing, and produces shareable scoring reports usable in CI pipelines.
open_source · generally available · Research snapshot 2026-09-06
Visit the official product source ↗Where it fits
Evaluation & testing · AI risk & compliance management
Useful conversation with: AI evaluation lead, Compliance manager, LLM application developer.
Ask for a demonstration
Show me running IMDA's starter-kit benchmarks plus an automated red-team attack module against our chatbot and the scoring report it produces.
Capabilities and evidence
Support labels reflect the supplied research. Documentation and vendor claims are not independent product tests. “Not established” means the researcher did not find support; it does not prove a capability is absent.
Documented by provider
The repository states Moonshot brings benchmarking and red-teaming together for LLM applications, tests bias, toxicity and hallucination alongside metrics such as accuracy and BLEU, supports guided workflows for IMDA's Starter Kit for LLM-based App Testing, and provides attack modules and custom recipes.
Limit: Benchmarks and attack modules are as good as their datasets; results are not certifications and dataset currency is not guaranteed.
Source s1
Documented by provider
AI Verify Foundation states Project Moonshot is an open-source LLM evaluation toolkit with a Python library and web UI, pre-built evaluators and benchmark datasets, CI/CD pipeline integration and shareable reports.
Limit: Foundation page is promotional in tone and does not quantify coverage or validity of the evaluators.
Source s2
Documented by provider
GitHub metadata records Apache-2.0 licensing, a non-archived repository and last push June 2026.
Limit: Activity gap of several months; maintenance cadence unclear.
Source s3
Limitations to discuss
- Evaluation coverage depends on bundled datasets
- No agent-specific (multi-step tool use) evaluation evidence found
Sources
- aiverify-foundation/moonshot · AI Verify Foundation / GitHub · official repository
Access date reported by researcher: 2026-09-06 - Project Moonshot · AI Verify Foundation · official product
Access date reported by researcher: 2026-09-06 - GitHub REST API repository record · GitHub · official repository
Access date reported by researcher: 2026-09-06
Listing does not imply partnership, supplier status, a working DutyGraph integration, or a compliance certification.
Suggest a correction