AI Plagiarism and Cheating Detection System for Education
With the increasing use of LLMs in student works, false positive rates of AI detectors reach 15%, leading to unfounded accusations and trust issues. Our combined approach reduces this to 2-4% using a three-level detection: stylometric analysis, MinHash for inter-student copying, and dynamic crawling of external sources. The model considers the context of the discipline: for humanities works the perplexity threshold is lower, for technical ones it is higher. Over 5 years we have implemented 20 systems in universities and EdTech projects, achieving up to 80% savings on work checking budgets. The system covers three scenarios: plagiarism from open sources (internet, dissertation databases), copying between students (with paraphrasing), and AI generation (GPT, Claude, LLaMA). Each scenario requires its own algorithm, so an ensemble of models provides the best accuracy. Developing the system pays off by reducing teacher workload and improving check quality.
How do we distinguish AI-generated text from human-written?
AI detectors (GPTZero, Originality.ai) often produce false positives on structured works with templated phrases. Our approach combines three signals:
- Text perplexity – LLM texts have anomalously low perplexity (each word is predictable). We use a GPT-2 model calibrated for the educational domain.
- Stylometric variability – sentence length, rare word frequency, conjunction usage. Students write unevenly, LLMs write uniformly.
- Semantic smoothness – overly coherent logical structure without speech errors. AI text lacks typical human 'noise'.
The joint model reduces the false positive rate by 60-70% compared to single detectors. More about perplexity on Wikipedia.
Why use MinHash for copying detection?
Full pairwise comparison of all works is O(n²). For 1000 works that would be 500,000 comparisons, taking hours. We use MinHash LSH – 10 times faster with 95% accuracy.
Click to view MinHash implementation
def detect_inter-student-similarity(submissions: list[Submission]) -> list[SimilarityPair]: # MinHash for approximate similarity from datasketch import MinHash, MinHashLSH lsh = MinHashLSH(threshold=0.4, num_perm=128) minhashes = {} for sub in submissions: m = MinHash(num_perm=128) for word in preprocess(sub.text).split(): m.update(word.encode('utf-8')) lsh.insert(sub.student_id, m) minhashes[sub.student_id] = m pairs = [] for sub in submissions: result = lsh.query(minhashes[sub.student_id]) for match_id in result: if match_id != sub.student_id: similarity = minhashes[sub.student_id].jaccard(minhashes[match_id]) if similarity > 0.4: pairs.append(SimilarityPair( student_1=sub.student_id, student_2=match_id, similarity=similarity )) return pairs The similarity threshold is adjustable: 0.4 for humanities, 0.6 for exact sciences. MinHash LSH is resilient to paraphrasing: at threshold 0.4 it catches 85% of paraphrased copies.
Handling Authorized Group Works
Group projects are a common cause of false positives. In our system, the administrator uploads a list of groups, and for students within the same group, similarity is not flagged as plagiarism. Additionally, a 'permitted borrowing' threshold is supported – e.g., up to 20% common phrases for technical disciplines. This is configurable in the interface.
What's Included in the Turnkey Solution
- ML Model (ensemble: stylometry + MinHash + AI detector)
- Web Interface for upload and reporting
- LMS Integration (REST API, plugins for Moodle/Blackboard)
- Comprehensive documentation (API spec, admin guide)
- Training sessions for teachers (2 sessions)
- 6 months of technical support post-deployment
Comparison of Detection Methods
| Method | Accuracy | Time for 1000 works | False Positive Rate |
|---|---|---|---|
| Simple MinHash | 93% | 30 sec | 8% |
| Semantic Comparison (BERT) | 97% | 5 min | 4% |
| Combined (Ours) | 97% | 2 min | 2-4% |
| AI Detector (GPTZero) | 85% | 10 sec | 15% |
The combined approach offers the best accuracy-speed ratio. Moreover, we conduct A/B testing on your dataset to calibrate thresholds. Our system is 2.5 times more accurate than GPTZero.
Our Experience and Guarantees
Certified ML engineers with experience in NLP and text processing. We have 5+ years on the market and have completed 20+ projects for universities and EdTech companies. Guarantees: false positive rate ≤ 5% on your dataset; integration with existing LMS in 2-3 days; 6 months of support post-deployment. For model quality monitoring we use Weights & Biases: after deployment, the model continues fine-tuning on new data, reducing concept drift. Implementation cost: typical projects range from $15,000 to $30,000 depending on scale, with an average savings of $50,000 per year for a university with 10,000 students.
Implementation Process
- Analysis – requirements gathering, process audit, data assessment.
- Design – model architecture, stack selection (PyTorch, Hugging Face, Vector DB).
- Development – model training, interface creation, integration.
- Testing – A/B test on real data, threshold calibration.
- Deployment – rollout, staff training, documentation handover.
Typical Implementation Mistakes
- Using a single AI detector as the sole evidence leads to false accusations.
- Ignoring authorized group works: the system must account for context.
- High similarity threshold (0.8) misses paraphrased copies.
Contact us for a preliminary evaluation of your project. Get an engineer consultation and a pilot launch in 4 weeks. Order a demo test on your data to verify the system's accuracy.







