Mulmocast is an open-source CLI tool that turns web articles, documents, papers and ideas into multi-modal content—like podcasts, slides, comics, and more—using the power of generative AIs.
Do you want to turn a web article into a comic strips? Do you wish there is a way to create a Youtube shorts video based on a scientific paper? MulmoCast is the answer.
Report
It's a really nice idea although it's implemented in TS. I've tried partially similar in python for my pitch slides / videos, not completely automated, since I was so tired of the repetitive efforts of that pitch slide work.
I'm wondering how to place the text on the right place in images?🤔 Or can gpt-imgae-1 generate with text embedded correctly as one-go, or with inpainting?
@hiroshi_doyu Texts within images are generated by gpt-image-1, and it is the result of the phrase "comic strips" in the image prompt. It is probably possible to specify the location and text in the prompt (in the generated script), but I'd rather automate it the whole process.
Report
Fits the current multimodal content trend beautifully! How's the output quality for different formats? Could be a game-changer for content repurposing 🎭
Report
Exciting to see a truly AI-native approach to presentations. Seamless integration of text, voice, and visuals opens up more dynamic and engaging communication.
Report
No reviews yetBe the first to leave a review for MulmoCast
InstaGraph
It's a really nice idea although it's implemented in TS. I've tried partially similar in python for my pitch slides / videos, not completely automated, since I was so tired of the repetitive efforts of that pitch slide work.
I'm wondering how to place the text on the right place in images?🤔 Or can gpt-imgae-1 generate with text embedded correctly as one-go, or with inpainting?
InstaGraph
@hiroshi_doyu Texts within images are generated by gpt-image-1, and it is the result of the phrase "comic strips" in the image prompt. It is probably possible to specify the location and text in the prompt (in the generated script), but I'd rather automate it the whole process.
Fits the current multimodal content trend beautifully! How's the output quality for different formats? Could be a game-changer for content repurposing 🎭
Exciting to see a truly AI-native approach to presentations. Seamless integration of text, voice, and visuals opens up more dynamic and engaging communication.