PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 24, 20260 citationsOpen Access

AI image generator – text to image and sketch to image generator

View Full Paper
RJRiya JadhavDLDr. Prashant Lokhande

Key Points

  • This research aims to develop an AI image generation platform that converts text and sketches into high-quality visuals.
  • Introduced a dual-model approach for image generation using text prompts and user drawings.
  • Utilized Google's Gemini multimodal API for generating images based on user inputs.
  • Employed Firebase for user authentication, cloud storage, and serverless computing.
  • Implemented a web interface where users can easily submit their inputs.
  • Demonstrated an average latency of under a few seconds for image generation.
  • The system effectively allows users to create images without requiring heavy GPU resources.
  • Proven capability to enhance creative workflows and enable real-time visualization.

Abstract

This research introduces an innovative AI image generation platform that utilizes a dual-model approach, capable of producing high-quality visuals from both text prompts and user-drawn sketches. The system's architecture cleverly pairs Google's Gemini multimodal API for its core generative power with the robust services of Firebase, handling user authentication, cloud storage, hosting, and all serverless backend tasks. Crucially, this design circumvents the need for heavy, GPU-intensive training typical of conventional AI, relying instead on API-based inference. This strategy effectively democratizes access to advanced image creation, making it available to students, educators, and developers without specialized hardware requirements. Users interact with a simple web interface, where their inputs are efficiently processed by Firebase Cloud Functions before being sent to Gemini. The final high-quality images are then instantly and securely stored in Firebase Storage and displayed. Demonstrating strong performance with average latency under a few seconds, the system proves that cloud-based multimodal AI can streamline creative workflows and enable real-time visualization with minimal infrastructure. Future developments aim to enhance style control, customization, and cross-platform capabilities.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Jadhav et al. (2026) studied this question.

synapsesocial.com/papers/69746149bb9d90c67120b2e3https://doi.org/10.5281/zenodo.18333916
Ask AI
Helpful
Bookmark
Share
View Full Paper