HyperAIHyperAI

Command Palette

Search for a command to run...

DocBank Text Dataset

Date

3 years ago

Size

48.1 GB

Organization

Beijing University of Aeronautics and Astronautics

Publish URL

github.com

Paper URL

arxiv.org

Featured Image

DocBank is a text dataset. The dataset contains 500,000 document pages with fine-grained, term-level annotations for document layout analysis. The dataset is constructed in a simple and effective way with weak supervision from \LaTeX{} documents available on arXiv.com.

DocBank.torrent
Seeding 2Downloading 0Completed 419Total Downloads 751
  • DocBank/
    • README.md
      967 字节
    • README.txt
      1.89 KB
      • data/
        • DocBank_500K_ori_img.zip.001
          5 GB
        • DocBank_500K_ori_img.zip.002
          10 GB
        • DocBank_500K_ori_img.zip.003
          15 GB
        • DocBank_500K_ori_img.zip.004
          20 GB
        • DocBank_500K_ori_img.zip.005
          25 GB
        • DocBank_500K_ori_img.zip.006
          30 GB
        • DocBank_500K_ori_img.zip.007
          35 GB
        • DocBank_500K_ori_img.zip.008
          40 GB
        • DocBank_500K_ori_img.zip.009
          45 GB
        • DocBank_500K_ori_img.zip.010
          47.41 GB
        • DocBank_500K_txt.zip
          47.9 GB
        • MSCOCO_Format_Annotation.zip
          48.1 GB

Build AI with AI

From idea to launch — accelerate your AI development with free AI co-coding, out-of-the-box environment and best price of GPUs.

AI Co-coding
Ready-to-use GPUs
Best Pricing
Get Started

Hyper Newsletters

Subscribe to our latest updates
We will deliver the latest updates of the week to your inbox at nine o'clock every Monday morning
Powered by MailChimp