Command Palette
Search for a command to run...
Hebrew-CLIP Hebrew Dataset
Hebrew-CLIP is a dataset released by NVIDIA in 2024 for training Hebrew visual-language models, aimed at promoting the application of visual-language models such as CLIP in Hebrew scenarios.
The dataset contains approximately 7.78 million Hebrew image descriptions, primarily composed of machine-translated data from DataComp-1B and multilingual data from LAION-5B. The data is stored in Parquet format and does not include actual images, only providing text descriptions and corresponding image embedding references, making it suitable for training and evaluating multimodal models.
Dataset Composition
The dataset consists of two Parquet files, each mainly containing the following fields:
- key: A unique identifier for the description.
- heb_caption: The Hebrew description text.
- file_name: The corresponding image embedding file name.
- file_index: The index position of the embedding within the file.
Build AI with AI
From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.