HyperAIHyperAI

Command Palette

Search for a command to run...

Cornell Researchers Develop WildFin Dataset for Fish Behavior Analysis

Researchers from Cornell University and partner institutions have introduced WildFin, a specialized dataset designed to advance computer vision capabilities for analyzing fish behavior in natural marine environments. The initiative addresses a critical gap in machine learning, where current models struggle to process wildlife footage due to a severe lack of scientific training data relative to internet-scale imagery. Abigail Grassick, a doctoral candidate in computational biology, led the creation of the dataset after observing that existing algorithms failed to track individual fish or identify distinct actions such as feeding or predator evasion during her fieldwork off the coast of Curaçao. WildFin comprises nine hours of frame-by-frame annotated underwater video captured from coral bommies. The collection process was resource-intensive, requiring approximately 1,400 hours of fieldwork and 600 hours of annotation to produce the final footage. The dataset includes diverse scenarios, ranging from complex scenes with roving fish bands to clips where divers tracked single individuals, the latter sourced from colleagues at the University of Colorado Boulder. These varied conditions were incorporated to test model robustness against challenges like variable lighting, shifting camera angles, and the difficulty of maintaining identity across frames. Despite the comprehensive annotation, attempts to fine-tune standard computer vision models on the WildFin data resulted in poor performance, underscoring the significant hurdles in adapting consumer-grade AI to ecological research. Andrew Hein, an associate professor of computational biology and study co-author, noted that the fraction of scientific imagery within general datasets is negligible. He emphasized that improved models could unlock the potential of citizen science and crowdsourced videos, enabling biologists to monitor ecosystem health and document rare species without the high costs associated with traditional expeditions. The project highlights the need for a multidisciplinary approach, combining expertise in statistics, machine learning, and ecology. Jennifer Sun, an assistant professor of computer science at Cornell, remarked that while computer vision has succeeded in controlled laboratory settings, wild environments present unique complexities that require tailored solutions. The team argues that addressing this gap demands collaboration across fields to make data acquisition less expensive and more scalable. Grassick and her colleagues presented their findings, titled WildFin: An In-the-Wild Dataset for Fish Behavioral Recognition, at the European Conference on Computer Vision in Malmö, Sweden, on September 8, with the work also available on the arXiv preprint server. The researchers issued a call to the broader scientific community, urging ecologists to release archived field videos stored on lab hard drives to serve as training data. Sun suggested that future efforts could involve artificial intelligence agents to automatically process and format these archives, effectively reviving dormant data for computational analysis. The team aims for WildFin to serve as a foundational resource, spurring the development of tools that allow researchers to interpret marine behavior with greater efficiency and precision.

Related Links