You open a porn generator, upload a photo, write what you want to get, and look at the result thinking: "I actually asked for something else."
A familiar situation.
Especially when you know exactly what you wanted to see. You already have an image in your head: a specific pose, clothing, gaze, lighting, angle. Perhaps you even imagine how the last second of the video should look.
But the AI doesn't see that picture in your head.
It only sees what is on the source photo and what you wrote.
Therefore, the problem often arises even before generation. You think in terms of a ready-made scene, but describe it with a few words. For you, "make it a bit bolder and more natural" might be a very specific request. For the model, it is not.
The good news is that you don't need to become a prompting specialist.
You just need to learn how to translate what you imagine into concrete visual details.
This article is exactly about that.
First, let's figure out what you are asking the AI for
SLYGEN has two different working scenarios.
If you are creating an image
You are essentially describing a future photograph.
You need to explain:
- who is in the frame;
- what the person looks like;
- what pose they are in;
- how they are positioned relative to the camera;
- what they are wearing;
- where they are located;
- what lighting and mood the scene has.
If you are creating a video from a photo
Most of this information is already there. The photo shows the person, their appearance, clothing, location, lighting, and initial body position. Therefore, your primary task is to tell what will happen next.
It turns out to be a simple scheme:
Image: what the frame should look like. Video: what should happen to the existing frame.
This is the first thing to remember. Further work with images and videos will differ slightly.
Part 1. How to get the desired image
Start not with "erotic," but with a specific task
The most common mistake is describing not the image, but the impression of it.
For example:
make the girl more attractive
or:
make the image more sensual
or:
let it be more revealing
For a human, everything is clear. But what exactly needs to change? Body position? Clothing? Gaze? Lighting? Angle? Cropping? For one person, a "sensual image" means bare shoulders and soft light. For another, it means a straight posture, looking at the camera, and a close-up.
The AI doesn't know this. Therefore, a general idea needs to be broken down into visible details.
Instead of:
make the image more sensual
you can specify:
bare shoulders, soft side lighting, looking straight into the camera, waist-up shot
Now the model has something to work with.
Don't explain what impression the frame should produce. Describe what that impression should consist of.
Build the image in layers
You don't have to write everything in one sentence.
It is more convenient to move from the main to the secondary:
character → composition → pose → clothing → environment → lighting → style
For example:
1 girl, long chestnut hair, blue eyes, evening dress, BREAK, full body, standing by the window, slightly turned towards the camera, one hand touching her hair, BREAK, modern bedroom, evening city outside the window, soft warm light, cinematic style
First, we defined the person.
Then, how they are positioned in the frame.
Then we added clothing and scene details.
This is easier to control than when everything is mixed into one long phrase.
Why is BREAK needed?
BREAK helps separate logical parts of the prompt.
For example:
who is in the frame BREAK pose and composition BREAK environment and lighting
But you don't need to use it after every word. If a part of the scene doesn't matter, just describe it briefly.
Don't try to write the prompt like a literary work
For a human, it sounds beautiful:
She seems to invite the viewer to linger on her, remaining in an atmosphere of light mystery...
But what exactly should the model do?
How is the person standing?
Where are they looking?
What are they wearing?
Where does the light fall from?
If there is a precise word for the needed detail, it is better to use it.
Especially be careful with:
- slang;
- metaphors;
- hints;
- ambiguous expressions;
- words whose meaning depends on context.
The phrase "makes it vulgar" works great with a human who truly understood. The AI cannot use this.
Therefore, the useful rule is simple:
Name the thing. Describe the position. Indicate the action.
Appearance, clothing, and pose are different tasks
Don't try to solve everything with one set of words at once.
If you like the character's appearance, you don't need to rewrite its description every time.
If the main problem is the pose, focus on the pose.
If you want to change the clothing, describe the clothing in detail.
This is especially noticeable in adult generation.
For example, "make the image more open" can be implemented in completely different ways. You can change the neckline, clothing length, shoulder position, silhouette, pose, or cropping.
Therefore, it is better to decide first what exactly needs to change.
Part 2. How to manage important details
Not all details are equally important
Imagine an image where the most important thing for you is the face and gaze.
Then it makes no sense to pay as much attention to the color of the wall in the background.
And vice versa: if the whole idea hinges on a specific pose, that is the one that needs to be described in more detail.
In SLYGEN, you can use weights from 0.1 to 2.0 for tags.
1.0 — normal influence.
More than 1.0 — strengthening.
Less than 1.0 — weakening.
For example:
(portrait:1.3), soft lighting, (blurred background:0.8)
Here the portrait becomes more significant, and the background less significant.
But weight does not mean "make better."
If you strengthen everything, priorities start to conflict.
Therefore, it is better to first get a normal scene without a large number of weights. Then strengthen only what really needs to become the main focus.
Change one thing at a time
Suppose the result is almost there, but you don't like the pose.
You don't need to change simultaneously:
- pose;
- clothing;
- lighting;
- background;
- angle;
- facial expression.
Otherwise, you won't understand what exactly helped.
It is simpler to change only the pose and look at the result.
Then, if necessary, the angle.
Then, lighting.
This way, a working combination is found gradually.
This is especially convenient when you already have a successful character and don't want to accidentally ruin everything else.
Part 3. How to work with photos and video
Now for the most important difference.
Video does not start from a blank sheet.
It starts with your photo.
The photo already sets the initial scene
If the photo shows a girl standing by a window in a specific dress, the video starts exactly from here.
Therefore, the request:
let her end up on the beach
does not solve the task of changing the location.
The window and room are still there in the photo.
If you need a different setting, first create an image with the desired scene. Then bring it to life.
The same goes for clothing.
If the source has an evening dress, it won't disappear just because you stopped mentioning it in the prompt.
Hence the main rule:
It is easier to change the scene with an image. Video is more convenient for developing an already ready-made scene.
Don't retell the photo
If the source image already has a person, a mirror, and a bedroom, you don't need to write again:
a girl with long hair stands in a bedroom near a mirror...
It is more important for you to tell what she is doing.
For example:
She slowly turns towards the camera, fixes her hair, and smiles.
Everything necessary to start the scene is already in the photo.
Part 4. How to describe movement correctly
One main action — for a short segment
In your head, you can easily imagine a whole scene:
she stands up, walks to the mirror, turns, fixes her dress, sits down, looks at the camera, and smiles.
But for a short video, this is too much.
A convenient benchmark:
5 seconds — one main action.
10 seconds — preparation and action.
15–20 seconds — preparation, action, and final state.
For example:
She gets out of bed and walks to the mirror. She stops, turns to the camera, and fixes her hair.
There is already a sequence here, but it is not overloaded.
If there are more actions, it is better to distribute them across several stages.
Start with the pose that already exists
The initial pose of the photo strongly influences what movement will look natural.
If a person is sitting, it is logical to continue the movement from a sitting position.
If standing — you can add a step, a turn, hand movement.
If lying down — start with head, hand, or torso movement.
Don't force the character to instantly change their entire body position without necessity.
For example, if a girl is sitting on the edge of the bed, it is much more natural to start with a turn towards the camera or a hand movement than to immediately send her across the room.
Don't start a new scene. Continue the one that is already in the photo.
A second person must appear in the prompt
If there is one person in the photo, and you write:
they look at each other
The AI might not understand who "they" are.
First, you need to introduce the second participant:
A man in a dark shirt approaches her.
After that, you can describe the joint action.
At the same time, it is important to consider the physics of the scene.
If characters are far apart, they cannot interact naturally without moving towards each other.
If a person needs to touch another, the body positions must allow this.
The more realistically you build the scene at the spatial level, the less you have to fix strange movements after generation.
Part 5. Camera and gaze
Angle is part of the concept
The same scene looks different depending on the camera position.
A close-up shows the face and gaze.
A waist-up shot allows you to see both facial expressions and the upper body.
Full body shows the pose and silhouette.
A side angle can better show movement.
Therefore, instead of:
make it beautiful
you can write:
camera is stationary
or:
camera slowly zooms in
or:
camera circles her from the left
or:
camera is at the side at chest level
or:
camera is at eye level
One precise indication is often enough.
Gaze is also better described concretely
"Make the gaze more seductive" sounds clear to a human but leaves the model a lot of freedom.
Much more specific:
looking straight into the lens
shifting gaze to the camera
looking over the shoulder
looking away to the side
closing eyes and smiling slightly
This can already be imagined as a specific action in the frame.
The same works with facial expressions. Instead of a general assessment, you can describe a smile, head position, gaze direction, or a small movement.
Part 6. Don't forget about sound
Video is not just an image.
If steps, rustling fabric, wind, rain, water, or the sound of a closing door are important, it is better to specify them.
The same goes for lines of dialogue.
For example:
She looks at the camera and says in Russian: "You didn't expect this."
For a short clip, it is better to use short lines — a benchmark is up to 15 words for five seconds.
And consider the physics: a person must have the opportunity to pronounce this phrase. If the face is turned away from the camera or the mouth is covered by an object, the result can be unpredictable.
Part 7. What to do if the AI still got it wrong
This is where many start rewriting the prompt from scratch.
Usually, this is unnecessary.
First, find one specific error.
Wrong pose
Check the body position and formulate it more precisely.
Wrong clothing
Check the source photo. Especially if it is about video.
Action is not readable
Look at the angle and the distance between objects.
The second person ended up in the wrong place
Check exactly where they appear and whether the body positions allow the action to be performed.
The last action disappeared
Most likely, you gave too many actions for this duration.
The pose turned out strange
Check if the body has support and if it can physically be in such a position.
Missing the required sound
Most likely, you just didn't specify it.
Instead of the question:
Why did the AI do nonsense again?
ask yourself another:
What specific information did I not give it?
This is much more productive.
And finally: don't look for a "secret prompt"
The most useful habit when working with SLYGEN is to stop looking for magic words.
Don't ask:
How to write the perfect prompt?
Ask:
What specifically needs to change on the screen?
Need to change the pose — describe the body position.
Need different clothing — name it.
Need a specific silhouette — think about the pose, clothing, and angle.
Need a different gaze — say where the person is looking.
A second character appears — introduce them into the scene.
An action happens — name it.
Need a specific ending — describe what should be visible at the end.
The camera needs to move — indicate the direction.
Need sound — name it.
That's it.
A good prompt doesn't have to look smart and doesn't have to impress the person reading it.
Its task is simpler: not to make the AI guess your idea.
You already know what picture you want to get.
Now you just need to learn how to explain it so that this picture is seen not only by you.
