kangarooking-skills  by kangarooking

AI Agent skills for multimodal content and workflow automation

Created 4 months ago
307 stars

Top 87.2% on SourcePulse

GitHubView on GitHub
Project Summary

This repository offers a curated collection of specialized AI agent "skills" designed to automate diverse tasks, ranging from multimedia processing and content generation to sophisticated workflow orchestration. It targets developers and power users building custom AI applications, providing modular, pre-built components that enhance agent capabilities and streamline complex operations.

How It Works

The project structures distinct functionalities into independent, self-contained "skills," each residing in its own subdirectory. This modular design facilitates easy selection, integration, and potential composition of skills within larger AI agent frameworks. Core technologies leveraged include APIs like APIMart GPT-Image-2 and Tencent Hunyuan, alongside tools such as LangChain, yt-dlp, and Whisper. This enables complex operations like multi-agent workflows, asynchronous task processing, and cross-platform content analysis and generation.

Quick Start & Requirements

  • Primary install / run command: Not explicitly detailed; likely involves cloning the repository and installing skill-specific Python dependencies.
  • Non-default prerequisites: API keys are frequently required (e.g., APIMart, twitterapi.io, Tencent Cloud SDK). Specific Python versions are not stated. Some skills may benefit from GPU acceleration (e.g., ASR, 3D generation).
  • Links: No direct links to official quick-start guides, demos, or comprehensive documentation are provided within the README.

Highlighted Details

  • apimart-image-gen: Asynchronous image generation with control over resolution (1k, 2k, 4k), aspect ratios, and API key management via environment variables.
  • harness-engineering: Implements a Plan-Build-Verify AI agent framework utilizing OpenAI Codex, Anthropic, and LangChain, featuring multiple specialized agents and workflow hooks.
  • video-downloader: Multi-platform video download (Douyin, Bilibili, YouTube, Xiaohongshu), original caption extraction, and ASR transcription using Whisper or SiliconFlow.
  • hy-3d-gen: Text-to-3D and image-to-3D model generation, compatible with TokenHub/OpenAI interfaces, supporting PBR materials and various generation modes.
  • viral-topic & viral-title: Advanced skills for cross-platform content discovery, topic ideation, and viral title generation, incorporating platform-specific methodologies and self-evolutionary title libraries.

Maintenance & Community

The repository welcomes contributions, with a defined structure for new skills (SKILL.md, references/). No specific community links (Discord, Slack) or details on maintainers, sponsorships, or roadmaps are provided.

Licensing & Compatibility

  • License type: MIT License.
  • Compatibility notes: The MIT license generally permits commercial use and integration into closed-source projects. However, users should verify the licenses of any third-party dependencies used by individual skills.

Limitations & Caveats

Several skills are described vaguely (e.g., reshape-your-life, task-harness) without explicit functionality details. Comprehensive setup instructions, specific dependency versions, and detailed usage examples beyond basic commands are not provided, potentially requiring users to infer or consult external documentation. The README does not explicitly mention testing procedures or performance benchmarks for the skills.

Health Check
Last Commit

5 days ago

Responsiveness

Inactive

Pull Requests (30d)
2
Issues (30d)
1
Star History
144 stars in the last 30 days

Explore Similar Projects

Feedback? Help us improve.