OʻZBEK TILIDAGI FIRIBGARLIK XABARLARINI ANIQLASHDA GIBRID YONDASHUV: QOIDALAR, PATTERNLAR VA ONNX NEYRON TARMOGʻINING BIRLASHUVI
Keywords:
firibgarlik aniqlash, fishing, smishing, oʻzbek tili, gibrid AI, qoidaga asoslangan tizim, pattern engine, neyron tarmoq, ONNX, TF-IDF, homoglyph hujumi, ijtimoiy muhandislik, RiskFusion, on-device AI, mobil xavfsizlik, matn klassifikatsiyasi.Abstract
Ushbu maqolada oʻzbek tilidagi SMS va messenjer xabarlari ichidan firibgarlik (phishing, smishing, scam) kontentini avtomatik aniqlash uchun gibrid yondashuv koʻrib chiqilgan. Tizim uchta mustaqil tahlil qatlamini birlashtiradi: 32 ta qoidaga asoslangan signal agregatsiyasi, 20 dan ortiq oldindan tuzilgan signal kombinatsiyalarini aniqlovchi pattern dvigateli (pattern engine) va 15 003 oʻlchovli TF-IDF asosidagi xususiyat vektori bilan ishlaydigan ONNX formatidagi neyron tarmoq modeli. Yakuniy xavf bahosi RiskFusion algoritmi orqali shakllantiriladi, u uchta manbani — Pattern Floor, Signal Risk va ML skorni — ustuvorlik boʻyicha birlashtiradi. Yondashuvning oʻziga xos jihati shundaki, butun tahlil mobil qurilmaning oʻzida (on-device) bajariladi va xabar matni hech qachon serverga uzatilmaydi. Bu oʻzbek tilidagi kontentni mahalliy modellar bilan tahlil qilish, kirill-lotin homoglyph hujumlariga qarshi himoya va past kechikishli (real-time) ishlash imkonini beradi.
References
1. Salton G., Buckley C. Term-weighting approaches in automatic text retrieval. // Information Processing & Management. — 1988. — Vol. 24, No. 5. — P. 513–523.
2. Manning C. D., Raghavan P., Schutze H. Introduction to Information Retrieval. — Cambridge University Press, 2008. — 482 p.
3. ONNX Runtime: cross-platform, high performance ML inferencing and training accelerator. — Microsoft Corporation, 2024. — URL: https://onnxruntime.ai/
4. Bai J., Lu F., Zhang K. et al. ONNX: Open Neural Network Exchange. — GitHub repository, 2019. — URL: https://github.com/onnx/onnx
5. Unicode Technical Standard #39: Unicode Security Mechanisms. — Unicode Consortium, 2023. — URL: https://www.unicode.org/reports/tr39/
6. Gabrilovich E., Gontmakher A. The homograph attack. // Communications of the ACM. — 2002. — Vol. 45, No. 2. — P. 128.
7. APWG Phishing Activity Trends Report. — Anti-Phishing Working Group, 2024. — URL: https://apwg.org/trendsreports/
8. Mishra S., Soni D. Smishing detector: A security model to detect smishing through SMS content analysis and URL behavior analysis. // Future Generation Computer Systems. — 2020. — Vol. 108. — P. 803–815.
9. URLhaus — abuse.ch threat intelligence platform. — abuse.ch, 2024. — URL: https://urlhaus.abuse.ch/
10. Google Safe Browsing API v4 Reference. — Google Developers, 2024. — URL: https://developers.google.com/safe-browsing/v4
11. VirusTotal Public API v3 Documentation. — VirusTotal (Google), 2024. — URL: https://docs.virustotal.com/
12. NIST Special Publication 800-63B: Digital Identity Guidelines — Authentication and Lifecycle Management. — National Institute of Standards and Technology, 2020.
13. Dworkin M. NIST Special Publication 800-38D: Recommendation for Block Cipher Modes of Operation: Galois/Counter Mode (GCM) and GMAC. — NIST, 2007.
14. Android Security Best Practices: Cryptography and the Keystore system. — Google Developers, 2024. — URL: https://developer.android.com/training/articles/keystore
15. Sanh V., Debut L., Chaumond J., Wolf T. DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. // arXiv preprint arXiv:1910.01108. — 2019.
16. McMahan H. B., Moore E., Ramage D. et al. Communication-Efficient Learning of Deep Networks from Decentralized Data. // Proceedings of AISTATS. — 2017. — P. 1273–1282.
17. MITRE ATT&CK Framework: Initial Access — Phishing (T1566). — The MITRE Corporation, 2024. — URL: https://attack.mitre.org/techniques/T1566/