Github Top Repositories
14.1K subscribers
2.82K photos
59 videos
10 files
2.94K links
Top GitHub repositories in one place 🚀
Explore the best projects in programming, AI, data science, and more.
Download Telegram
Github Top Repositories
Photo
📌 Spotted on GitHub Trending: huggingface/speech-to-speech — let's break it down.

🔗 https://github.com/huggingface/speech-to-speech
📝 Build local voice agents with open-source models
──────────────────────────────

The huggingface/speech-to-speech repository offers a fully modular voice-agent pipeline, including Voice Activity Detection (VAD), Speech to Text (STT), Language Model (LLM), and Text to Speech (TTS) capabilities. This pipeline is exposed through an OpenAI Realtime-compatible WebSocket API, allowing for low-latency and flexible integration. Every component in the pipeline is swappable, enabling users to choose from a variety of backends and models, such as Parakeet TDT for STT, OpenAI-compatible LLM, and Qwen3-TTS for speech output.

To get started, you can install the package using pip install speech-to-speech and then run the server with speech-to-speech. The repository also includes a Quickstart section with example commands to demonstrate its usage. For instance, you can use the python scripts/listen_and_play_realtime.py --host 127.0.0.1 --port 8765 command to talk to the server from a second terminal.

The pipeline is designed to be highly customizable, with support for various backends and models. You can select specific implementations using CLI flags, such as --stt, --llm_backend, and --tts. The repository also includes a list of supported components, including VAD, STT, LLM, and TTS, along with their corresponding backends and platforms.

The target audience for this repository includes developers and researchers interested in building voice agents and exploring the capabilities of speech-to-speech models. With its modular design and flexible API, the huggingface/speech-to-speech repository provides a powerful tool for creating custom voice agents and integrating them into various applications.

In summary, the huggingface/speech-to-speech repository is a versatile and highly customizable pipeline for building voice agents, offering a wide range of backends and models to choose from. Whether you're a developer or researcher, this repository provides a valuable resource for exploring the possibilities of speech-to-speech technology.
Takeaway: Build your own voice assistant with huggingface/speech-to-speech and unlock a world of possibilities in speech-to-speech technology.

──────────────────────────────
🧠 Channel: https://t.iss.one/GithubRe
abus-aikorea/voice-pro is making waves. Here's the full picture.

🔗 https://github.com/abus-aikorea/voice-pro
📝 Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
──────────────────────────────

Voice-Pro is an AI-powered web application designed for speech recognition, translation, and multilingual dubbing. Its key features include top-tier speech recognition with Whisper, Faster-Whisper, and Whisper-Timestamped, as well as zero-shot voice cloning with F5-TTS, E2-TTS, and CosyVoice. The application also supports multilingual text-to-speech with Edge-TTS and kokoro, and offers instant translation for over 100 languages.

To use Voice-Pro, simply download and install the application, then access its web interface to upload your media files and select the desired features. The application is designed to be user-friendly and accessible, with a clean and intuitive interface.

From a technical standpoint, Voice-Pro is built using a range of cutting-edge technologies, including Python 3.12, Torch 2.8.0+cu128, and Gradio 6.20. It also utilizes several AI models, including Whisper, F5-TTS, and CosyVoice, to deliver high-quality speech recognition and voice cloning capabilities.

Voice-Pro is designed for a range of users, including creators, researchers, and multilingual professionals. It offers a robust alternative to ElevenLabs, with advanced voice solutions and a user-friendly interface.

In short, Voice-Pro is a powerful tool for anyone looking to leverage the power of AI for speech recognition, translation, and multilingual dubbing - and with its open-source code and free distribution, the possibilities are endless!

──────────────────────────────
🧠 Channel: https://t.iss.one/GithubRe
Github Top Repositories
Photo
🎯 iv-org/invidious landed on trending. Worth a proper look.

🔗 https://github.com/iv-org/invidious
📝 Invidious is an alternative front-end to YouTube
──────────────────────────────

Invidious is an open-source alternative front-end to YouTube, offering a lightweight and ad-free experience. Its key features include no tracking, customizable homepage, and audio-only mode. Users can import subscriptions from YouTube and other platforms, and export them as needed.

From a technical standpoint, Invidious has an embedded video support and a developer API, making it a great option for developers. The project is hosted on GitHub and has a large community of contributors, with translations available in many languages.

To get started, users can select a public instance or host Invidious themselves by following the installation instructions. The project is suitable for anyone looking for a private and customizable YouTube experience.

Invidious is perfect for those who want to ditch YouTube's ads and tracking - and it's completely free and open-source. Join the Invidious community today and experience the power of open-source video sharing!

──────────────────────────────
🧠 Channel: https://t.iss.one/GithubRe
🌟 ansible/ansible caught my eye on GitHub Trending today.

🔗 https://github.com/ansible/ansible
📝 Ansible is a radically simple IT automation platform that makes your applications and systems easier to deploy and maintain. Automate everything from code deployment to network configuration to cloud management, in a language that approaches plain English, using SSH, with no agents to install on remote systems.https://docs.ansible.com.
──────────────────────────────

Ansible is a simple IT automation system that handles configuration management, application deployment, and more. Its key features include an extremely simple setup process, agentless architecture, and the ability to describe infrastructure in a human-friendly language.

To use Ansible, you can install it with pip or a package manager, and power users can run the devel branch for the latest features. The community is active, with a forum for asking questions, getting help, and interacting with other users.

From a technical perspective, Ansible focuses on security and auditability, and allows module development in any dynamic language. The project is coded in a variety of languages, including Python, and has a devel branch for ongoing development.

Ansible is suitable for a wide range of users, from system administrators to developers, and is widely used in the industry.

The project is licensed under the GNU General Public License v3.0 or later.

In short, Ansible is all about making complex IT tasks radically simple - and that's a game-changer!

──────────────────────────────
🧠 Channel: https://t.iss.one/GithubRe
💡 microsoft/TRELLIS.2 just hit the trending charts — here's why it matters.

🔗 https://github.com/microsoft/TRELLIS.2
📝 Native and Compact Structured Latents for 3D Generation
──────────────────────────────

TRELLIS.2 is a state-of-the-art 3D generative model that enables high-fidelity image-to-3D generation. It features a novel "field-free" sparse voxel structure called O-Voxel, which allows for the reconstruction and generation of complex 3D assets with sharp features and full PBR materials. The model boasts high-quality, resolution, and efficiency, generating high-resolution fully textured assets with exceptional fidelity.

Some of the key features of TRELLIS.2 include:
- Arbitrary topology handling: The O-Voxel representation can handle complex structures without lossy conversion, including open surfaces, non-manifold geometry, and internal enclosed structures.
- Rich texture modeling: The model supports arbitrary surface attributes, including base color, roughness, metallic, and opacity, enabling photorealistic rendering and transparency support.
- Minimalist processing: Data processing is streamlined for instant conversions that are fully rendering-free and optimization-free.

To use TRELLIS.2, you'll need to install the dependencies, including the CUDA Toolkit and Conda, and then clone the repository. You can then use the pretrained model for image-to-3D generation or PBR texture generation. The repository also provides a web demo for easy testing.

From a technical standpoint, TRELLIS.2 is a 4B-parameter model that utilizes a Sparse 3D VAE with 16× spatial downsampling to encode assets into a compact latent space. The model is designed for high-performance and can generate high-resolution assets with exceptional fidelity.

The target audience for TRELLIS.2 includes researchers, developers, and artists working with 3D generation and image-to-3D applications. With its powerful features and streamlined processing, TRELLIS.2 is an ideal choice for anyone looking to push the boundaries of 3D generation.

In short, TRELLIS.2 is a game-changer for 3D generation, offering unparalleled fidelity, efficiency, and flexibility - and with its open-source availability, the possibilities are endless!

──────────────────────────────
🧠 Channel: https://t.iss.one/GithubRe