feat: use raw pdf inline data for multimodal gemini extraction of sample images
Build and Deploy / build-and-push (push) Successful in 38s

This commit is contained in:
AI Bot
2026-09-02 13:03:35 +05:30
parent 253c705f52
commit 1a99d451a7
+16 -8
View File
@@ -333,22 +333,30 @@ app.post('/api/sparkplug/extract', authenticate, upload.single('file'), async (r
// Parse PDF
const dataBuffer = fs.readFileSync(req.file.path);
const pdfData = await pdfParse(dataBuffer);
const textContent = pdfData.text;
const base64Pdf = dataBuffer.toString('base64');
// Call Gemini to extract prompt
const genAI = new GoogleGenerativeAI(geminiKey);
const model = genAI.getGenerativeModel({ model: "gemini-3.5-flash-lite" });
const prompt = `
You are an expert product designer. Read the following text from a product spec PDF and create a highly detailed, concise visual design prompt for an image generation AI.
Focus on the object's shape, color, materials, packaging details, and label content. Do not include background details.
TEXT:
${textContent.substring(0, 30000)} // Limit to avoid context length issues if massive
You are an expert product designer. Review the attached product spec PDF (which may include text and sample images).
Your goal is to create a highly detailed, concise visual design prompt for an image generation AI.
CRITICAL: If there is a sample image of the product in the PDF, you MUST carefully analyze it and extract the exact HEX color codes used in the design.
Include these HEX codes, as well as the exact shape, materials, typography style, and label content in your final prompt so the image generator knows exactly how it looks.
Do not include background details.
`;
const result = await model.generateContent(prompt);
const result = await model.generateContent([
{
inlineData: {
data: base64Pdf,
mimeType: "application/pdf"
}
},
prompt
]);
const response = await result.response;
const extractedPrompt = response.text();