# Max\_Molecule\_Size Error

**URL:** <https://ask.bioexcel.eu/t/max-molecule-size-error/4821>\
**Category:** HADDOCK\
**Created:** [February 12, 2024, 6:10am UTC](https://ask.bioexcel.eu/t/max-molecule-size-error/4821 "2024-02-12T06:10:43Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Supantha](https://avatars.discourse-cdn.com/v4/letter/s/ec9cab/32.png) [@Supantha](https://ask.bioexcel.eu/u/Supantha)\
**Post date:** [February 12, 2024, 6:10am UTC](https://ask.bioexcel.eu/t/max-molecule-size-error/4821/1 "2024-02-12T06:10:43Z")

</div>

Hello,  
I want to get the score of around 600 docked files. If I input 15/20 files, the scoring gives successful output. However, when I input all the molecule in molecules=[" "] section, it gives following error:  
2024-02-12 00:06:46,602 libutil ERROR] Too many molecules defined, max is {max\_molecules\_allowed}.  
Traceback (most recent call last):  
File “/opt/haddock3/src/haddock/libs/libutil.py”, line 310, in log\_error\_and\_exit  
yield  
File “/opt/haddock3/src/haddock/clis/cli.py”, line 135, in main  
params, other\_params = setup\_run(  
File “/opt/haddock3/src/haddock/gear/prepare\_run.py”, line 262, in setup\_run  
raise ConfigurationError(“Too many molecules defined, max is {max\_molecules\_allowed}.”) # noqa: E501  
haddock.core.exceptions.ConfigurationError: Too many molecules defined, max is {max\_molecules\_allowed}.  
[2024-02-12 00:06:46,603 libutil ERROR] Too many molecules defined, max is {max\_molecules\_allowed}.  
[2024-02-12 00:06:46,603 libutil ERROR] An error has occurred, see log file. And contact the developers if needed.  
[2024-02-12 00:06:46,603 libutil INFO] Finished at 12/02/2024 00:06:46. For any help contact us at [Issues · haddocking/haddock3 · GitHub](https://github.com/haddocking/haddock3/issues). Adéu-siau! Ciao! Au revoir!.

What should I do in this situation if I want to get the score from those 600 complexes.

Thank you!

---

<div class="post-metadata">

**Author:** ![Supantha](https://avatars.discourse-cdn.com/v4/letter/s/ec9cab/32.png) [@Supantha](https://ask.bioexcel.eu/u/Supantha)\
**Post date:** [February 12, 2024, 6:16am UTC](https://ask.bioexcel.eu/t/max-molecule-size-error/4821/2 "2024-02-12T06:16:17Z")

</div>

I want to add that this is Haddock3, and I am running it in cluster.

---

<div class="post-metadata">

**Author:** ![marco.giulini](https://dub1.discourse-cdn.com/flex013/user_avatar/ask.bioexcel.eu/marco.giulini/32/961_2.png) [@marco.giulini](https://ask.bioexcel.eu/u/marco.giulini)\
**Post date:** [February 12, 2024, 8:41am UTC](https://ask.bioexcel.eu/t/max-molecule-size-error/4821/3 "2024-02-12T08:41:58Z")

</div>

Hi,

you should merge all the molecules together in an ensemble pdb file, as in [haddock3/examples/scoring/capri-scoring-test.cfg at main · haddocking/haddock3 · GitHub](http://github.com/haddocking/haddock3/blob/main/examples/scoring/capri-scoring-test.cfg) . For that you can use the following command:

`pdb_mkensemble model_1.pdb ... model_600.pdb | pdb_tidy > my_ensemble.pdb`

This uses the pdbtools package that is part of the haddock3 python environment.

---

<div class="post-metadata">

**Author:** ![amjjbonvin](https://dub1.discourse-cdn.com/flex013/user_avatar/ask.bioexcel.eu/amjjbonvin/32/23_2.png) [@amjjbonvin](https://ask.bioexcel.eu/u/amjjbonvin)\
**Post date:** [February 12, 2024, 8:48am UTC](https://ask.bioexcel.eu/t/max-molecule-size-error/4821/4 "2024-02-12T08:48:02Z")

</div>

What you can do is to create an ensemble PDB file with all your models.

This can be done for example with the pdb\_mkensemble command from pdb-tools

You then only specify one input file in the scoring script

Are you using the emscoring.cfg example from the examples/scoring directory?

---

<div class="post-metadata">

**Author:** ![Supantha](https://avatars.discourse-cdn.com/v4/letter/s/ec9cab/32.png) [@Supantha](https://ask.bioexcel.eu/u/Supantha)\
**Post date:** [February 12, 2024, 7:09pm UTC](https://ask.bioexcel.eu/t/max-molecule-size-error/4821/5 "2024-02-12T19:09:25Z")

</div>

Hi,  
Thank you both. pdb\_mkensemble is working.

And yes, I am following the emscoring.cfg example. However, I have a question. Should I also include md\_scoring? The purpose is just to get the score of all complexes and get the best ones.

(script:  
directory in which the scoring will be done  
run\_dir = “/work/…”

# execution mode

ncores = 64  
mode = “local”

# ensemble to be scored

# Directory containing the ClusPro folder

molecules=[“my\_ensemble.pdb”]

# ====================================================================

# Parameters for each stage are defined below

[topoaa]

[emscoring]  
tolerance = 20

[caprieval]

---

<div class="post-metadata">

**Author:** ![amjjbonvin](https://dub1.discourse-cdn.com/flex013/user_avatar/ask.bioexcel.eu/amjjbonvin/32/23_2.png) [@amjjbonvin](https://ask.bioexcel.eu/u/amjjbonvin)\
**Post date:** [February 12, 2024, 10:20pm UTC](https://ask.bioexcel.eu/t/max-molecule-size-error/4821/6 "2024-02-12T22:20:15Z")

</div>

emscoring should be fine (and much faster)

If you have more clashes though, you could consider mdscoring, but it will take more time as it performs a very short MD simulation in explicit water.
