Artificial intelligence is increasingly used to support natural product identification, bioactivity prediction, target inference, molecular property estimation, and candidate prioritization, yet model evaluation frequently remains dominated by aggregate predictive performance. Such evaluation can obscure structural errors, unreliable annotations, data leakage, narrow chemical-space coverage, distribution shifts, mechanistic uncertainty, and limited relevance to experimental decision-making. This article proposes a conceptual benchmarking framework for AI in natural product drug discovery that organizes evaluation around four interdependent dimensions: data integrity, model robustness, mechanistic validity, and translational utility. The framework begins with natural product identity, chemical structure, stereochemistry, provenance, bioactivity, assay, target, and safety-data audits. It then requires task-aligned split strategies, transparent baseline comparisons, external and out-of-distribution testing, perturbation analysis, uncertainty estimation, bias assessment, and reproducibility documentation. Mechanistic validity is treated as an evidence-integration problem involving compound–target support, biological context, pathway plausibility, omics consistency, assay relevance, exposure considerations, toxicity, and experimental perturbation. Translational utility is evaluated through experimental actionability, safety awareness, developability, source authenticity, external generalizability, implementation feasibility, and expert review. The proposed architecture establishes decision boundaries that prevent benchmark success from being interpreted as proof of mechanism, therapeutic efficacy, clinical utility, safety, deployment readiness, or regulatory acceptance. Its principal contribution is a structured evaluation logic intended to strengthen model comparison, expose limitations, guide research prioritization, and improve the credibility and reproducibility of AI-assisted natural product discovery.