Discover and explore top open-source AI tools and projects—updated daily.
yueduanFine-grained binary diffing tool for x86 binaries
Top 99.9% on SourcePulse
Summary
DeepBinDiff is a fine-grained binary diffing tool specifically engineered for x86 architectures. It tackles the complex challenge of identifying precise code differences between two binary executables at a granular level. Primarily targeting security researchers, reverse engineers, and developers involved in binary analysis, DeepBinDiff offers a systematic method for comparing basic blocks. This capability is crucial for tasks such as vulnerability discovery, understanding software evolution, and detecting code plagiarism.
How It Works
The tool's architecture is built upon the angr framework, a popular framework for binary analysis. Its distinctive approach involves an on-the-fly Natural Language Processing (NLP) training mechanism. This process uniquely utilizes only the two input binaries for training, thereby circumventing the common Out-Of-Vocabulary (OOV) problem often encountered with pre-trained models. While this method enhances robustness by avoiding OOV issues, it is noted to potentially increase the overall processing time required for diffing.
Quick Start & Requirements
python3 src/deepbindiff.py --input1 <path_to_first_binary> --input2 <path_to_second_binary> --outputDir <output_directory>.tensorflow (version >= 1.14.0 and < 2.0), gensim, angr, networkx, lapjv, and a python3 environment.Highlighted Details
src/analysis_in_batch.sh, facilitating batch processing of multiple binary comparisons.output/nodeIndexToCode file.Maintenance & Community
The provided README text does not contain information regarding project maintainers, community support channels (such as Discord or Slack), active development status, or notable contributors.
Licensing & Compatibility
No specific licensing details or compatibility notes for commercial use or integration with closed-source projects were mentioned in the provided README text.
Limitations & Caveats
The tool's current implementation is restricted to CPU processing, lacking GPU acceleration which could significantly impact performance on large datasets. The on-the-fly NLP training, while addressing OOV challenges, inherently leads to longer execution times compared to approaches that leverage pre-existing, trained models.
4 years ago
Inactive
bigcode-project
src-d