AI News HubLIVE
站内改写3 分钟阅读

待翻译:Dataset: AI agent security failures, 1000 incidents classified

AI 服务暂时不可用,以下为来源摘要,待恢复后补全翻译:...\n **config_kwargs,\n )\n"," File \"/usr/local/lib/python3.14/site-packages/datasets/inspect.py\", line 291, in get_dataset_config_info\n raise SplitsNotFoundError(\"The split names could not be parsed from the datas…

来源Hacker News AI作者: LegionAPI

AI 服务暂时不可用,以下为来源正文,待恢复后补全翻译。

...\n config_kwargs,\n )\n"," File \"/usr/local/lib/python3.14/site-packages/datasets/inspect.py\", line 291, in get_dataset_config_info\n raise SplitsNotFoundError(\"The split names could not be parsed from the dataset config.\") from err\n","datasets.inspect.SplitsNotFoundError: The split names could not be parsed from the dataset config.\n"]},"partial":false,"jwt":"eyJhbGciOiJFZERTQSIsImtpZCI6IjVHZDBvd0g5MTM2eDZjc1FvbE1zcktNWUZoRVFoUm5PVVVybEpjOUhaUEEifQ.eyJyZWFkIjp0cnVlLCJwZXJtaXNzaW9ucyI6eyJyZXBvLmNvbnRlbnQucmVhZCI6dHJ1ZX0sImlhdCI6MTc4Nzc0MDE0MywianRpIjoiNGZiYzQ0MGMtNTdhNS00Y2M1LWE3MWItZjUxNDg2NDI5NTMwIiwic3ViIjoiL2RhdGFzZXRzL2dlbW1vemVyby9haS1hZ2VudC1zZWN1cml0eS1pbmNpZGVudHMiLCJleHAiOjE3ODc3NDM3NDMsImlzcyI6Imh0dHBzOi8vaHVnZ2luZ2ZhY2UuY28ifQ.v67oZlfiDt5W3XfFG-9i2klqqRlesaE7EOj2bAhQ8zfU9dajri0fISmxPDkgTJVipagRTW3zN1PEQbDJ1hodCg","displayUrls":true},"dataset":"gemmozero/ai-agent-security-incidents","isGated":false,"isPrivate":false,"hasParquetFormat":false,"isTracesDataset":false,"author":{"_id":"69e59c398fef259c9600110c","avatarUrl":"/avatars/ed6df6a44b34bcf83fe9146edc74ef63.svg","fullname":"Gemmo Zero","name":"gemmozero","type":"user","isPro":false,"isHf":false,"isHfAdmin":false,"isMod":false,"isUserFollowing":false},"compact":true,"isLoggedIn":false}"> Dataset Viewer The dataset viewer is not available for this subset. Cannot get the split names for the config 'default' of the dataset. Exception: SplitsNotFoundError Message: The split names could not be parsed from the dataset config. Traceback: Traceback (most recent call last): File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 290, in _generate_tables pa_table = paj.read_json( io.BytesIO(batch), read_options=paj.ReadOptions(block_size=block_size) ) File "pyarrow/_json.pyx", line 342, in pyarrow._json.read_json File "pyarrow/error.pxi", line 155, in pyarrow.lib.pyarrow_internal_check_status File "pyarrow/error.pxi", line 92, in pyarrow.lib.check_status raise convert_status(status) pyarrow.lib.ArrowInvalid: JSON parse error: Column() changed from object to string in row 0 During handling of the above exception, another exception occurred: Traceback (most recent call last): File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 286, in get_dataset_config_info for split_generator in builder._split_generators( ~~~~~~~~~~~~~~~~~~~~~~~~~^ StreamingDownloadManager(base_path=builder.base_path, download_config=download_config) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ ) ^ File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 101, in _split_generators pa_table = next(iter(self._generate_tables(splits[0].gen_kwargs, allow_full_read=False)))[1] ~~~~^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/usr/local/lib/python3.14/site-packages/datasets/packaged_modules/json/json.py", line 304, in _generate_tables batch = json_encode_fields_in_json_lines(original_batch, json_field_paths) File "/usr/local/lib/python3.14/site-packages/datasets/utils/json.py", line 111, in json_encode_fields_in_json_lines examples = [ujson_loads(line) for line in original_batch.splitlines()] ~~~~~~~~~~~^^^^^^ File "/usr/local/lib/python3.14/site-packages/datasets/utils/json.py", line 20, in ujson_loads return pd.io.json.ujson_loads(*args, kwargs) ~~~~~~~~~~~~~~~~~~~~~~^^^^^^^^^^^^^^^^^ ValueError: Expected object or value The above exception was the direct cause of the following exception: Traceback (most recent call last): File "/src/services/worker/src/worker/job_runners/config/split_names.py", line 68, in compute_split_names_from_streaming_response for split in get_dataset_split_names( ~~~~~~~~~~~~~~~~~~~~~~~^ path=dataset, ^^^^^^^^^^^^^ config_name=config, ^^^^^^^^^^^^^^^^^^^ token=hf_token, ^^^^^^^^^^^^^^^ ) ^ File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 340, in get_dataset_split_names info = get_dataset_config_info( path, ...... config_kwargs, ) File "/usr/local/lib/python3.14/site-packages/datasets/inspect.py", line 291, in get_dataset_config_info raise SplitsNotFoundError("The split names could not be parsed from the dataset config.") from err datasets.inspect.SplitsNotFoundError: The split names could not be parsed from the dataset config. Need help to make the dataset viewer work? Make sure to review how to configure the dataset viewer, and open a discussion for direct support. AI Agent Security Incidents 2026 1,000+ classified incidents • Updated daily • CC BY 4.0 Pipeline entièrement local (Lenovo Legion, CPU only, Ollama + Llama 3.1 8B). Collecte automatique depuis NVD/CVE, GitHub Advisories, Hacker News et de multiples sources de sécurité. Statistiques principales : 1 000+ incidents acceptés 82 incidents critiques Top vendors : NVIDIA (48), OpenAI (46), TensorFlow (44) Top types : api_exploit, unauthorized_action, configuration_exploit, data_exfiltration, prompt_injection, sandbox_escape, slopsploit_attack_chain (découvert par le modèle) Toutes les entrées ont une source_url vérifiable. Dataset complet : https://huggingface.co/datasets/gemmozero/ai-agent-security-incidents Code du pipeline : https://github.com/Legion33shadow/legion-n8n-shield Feedback technique bienvenu (faux positifs, taxonomy, nouvelles sources). Dernière mise à jour : 23 août 2026 Support this project If you find this dataset useful, consider supporting continued development: https://gemmo.gumroad.com/l/mdevxu Downloads last month 178 Total file size: 2.69 MB