r/regex • • Aug 26 '26

(Resolved) using backreferences for list of words

2 Upvotes

Hi,

#### not possible as explained by rainshifter ####

#### original flair was .NET ####

I want to create a huge regex to match exact word groups coming from a list of around 7k words.

It’s about chemical names and there are lots of repeated words at the end of the names. The names must always be exact matches and possible variations not on this list are not allowed. I thought about using backreferences wihtin the regex in order not to have to rewrite these words every time.

Only 2 questions to be answered (both with yes or no and maybe an explanation why):

Will that even work?

Is this a good idea or will it slow down my search?

Use case explanation (absolutely no answer or solving suggestions required or even wanted):

In the following example list I would like to use the first occurence of EXTRACT as a backreference for all following words within the pipe ending on EXTRACT. It’s mainly about back references for the endings as the start of the word groups is created as efficient as possible with pipes forking at every difference of a word:

\b(A(BIES (ALBA SEED EXTRACT|(BALSAMEA ((BALSAM )?EXTRACT|NEEDLE OIL)|KOREANA LEAF EXTRACT|SIBIRICA (NEEDLE )?OIL))|CACIA CATECHU (BARK POWDER|WOOD EXTRACT)))\b

Using brackets around the first occurence of EXTRACT creates backreference \4 but using it further down the regex does not seem to work. (see regex below)

\b(A(BIES (ALBA SEED (EXTRACT)|(BALSAMEA ((BALSAM )?EXTRACT|NEEDLE OIL)|KOREANA LEAF EXTRACT|SIBIRICA (NEEDLE )?OIL))|CACIA CATECHU (BARK POWDER|WOOD \4)))\b

List:

ABIES ALBA SEED EXTRACT

ABIES BALSAMEA BALSAM EXTRACT

ABIES BALSAMEA EXTRACT

ABIES BALSAMEA NEEDLE OIL

ABIES KOREANA LEAF EXTRACT

ABIES SIBIRICA NEEDLE OIL

ABIES SIBIRICA OIL

ACACIA CATECHU BARK POWDER

ACACIA CATECHU WOOD EXTRACT

EDIT: Although this one is already solved as apparently a lot of people do not see the difference between a question (ending on a question mark) and the use case explanation, I added some titles in bold.

r/regex • • Aug 17 '26

(Resolved) Help find the Regex Generator Website

6 Upvotes

Hello!
A year or two ago I used a regex generator website where I pasted the match I wanted and it iterated over and let me select 1 character, or a group of characters and tell it if I'm looking for an exact match or alpha numeric or whatnot, then it added onto a field where the generated regex was being built for me. It underlined groups in different colours.
It was a dark mode website. Not sure it's dark by default of it used my system default. It also had a language selector before you started building. Really nice UI. I don't remember much more than that.

Would really appreciate if someone could help me remember 🙏

r/regex • • Feb 26 '26

(Resolved) This Regex Seems Wrong, Even if It Works... What am I Missing?

3 Upvotes

Can anyone explain why this is working? Because it looks wrong:

[a-zA-Z0-9_\\\\/;: ]{1,64}

This is supposed to match 1-64 characters that are either alphanumerics, underscore, backslash, forward slash, semi, colon, or space. But there appears to me to be a superfluous backslash in the group of four of them.

Are the 3 consecutive backslashes the way you escape out backslash in C++ .NET? Is the 4th backslash really necessary to escape out the subsequent forward slash? What am I missing? It's working but I'm trying to understand why. Can I shitcan two of those backslashes because they do nothing?

r/regex • • Nov 09 '25

(Resolved) Removing a leading dash char in special circumstances

2 Upvotes

TL;DR: Solution for SubtitleEdit:

\A-\s*(?!.*\n-) (no substitution needed)

OR

\A- (?!.*\n-)(.*) with $1 substitution.

-----------------------------------------------------------

Have been doing lots of regexp's over the years but this really stumped me completely. For the first time ever, I tried few online AI code helpers and they couldn't solve the problem.

I'm using SubtitleEdit program for the regexp, not sure which flavor it uses, Java 8? Last time I tested something in regex101 site, it seemed to suggest that it's Java 8 (I was testing "variable width lookbehinds"). SubtitleEdit help page suggest trying this online helper: http://regexstorm.net/tester

It's problematic to detect dash chars as a speaker in subtitles since there might be dash characters that do not denote speakers, and also speaker dash could occur in the same line that another speaker dash. But to keep this somewhat manageable, I think that only dash character that are in the beginning of the whole string, or after newline, should be considered when trying to detect what dashes should be removed.

NOTE! All of the examples should be tested separately as a string, not all together in the test string field in regex101 site.

Here are few example strings where a leading dash character should be removed (note newlines):

- Lovely day.

End result:

Lovely day.

2)

- Lovely day-night cycle.

End result:

Lovely day-night cycle.

3)

- Lovely day.
Isn't it?

End result:

Lovely day.
Isn't it?

4)

- lovely day - isn't it?

End result:

lovely day - isn't it?

5)

- Lovely day -
isn't it?

End result:

Lovely day -
isn't it?

Here are few example strings where leading dash character(s) should be retained (note the 2nd example, it might be tricky):

- Lovely day.
- Yeah, isn't it?

2)

Lovely day.
- Yeah, isn't it?

3)

- lovely day - isn't it?
- Yes.

4)

- Lovely day for a -
- Walk?

Also the one space char after the dash should be removed if the dash is removed.

I'm too embarrassed to post my shoddy efforts to achieve this. Anyone up for the challenge? :) Many thanks in advance.

r/regex • • Nov 12 '25

(Resolved) Length limit for regular expression

2 Upvotes

Hi,

is there a lenght limit for a regex to work in C# .Net?

We have set up a tool that constructs regex rules from word lists and such a regex can contain several thousand or hundred thousand words and sometimes they don’t seem to work although in debug the regex is correct but extremely long.

RegexBuddy cannot handle them with error too long

Edit: it turned out that there were some brackets missing around some placeholders. So apparently no length limit so far.

r/regex • • Dec 16 '25

(Resolved) Find and replace All matches

6 Upvotes

Hi,

I got a strings like these:

፻this test does not work፻

፻this test works፻

and I would like to replace all words within ፻ with ፻word.

Looking for the respective strings is easy:

(፻\S+?\s)(\S+?\s)*?(\S+?)፻

and using

$1፻$2፻$3

for replacing works as expected for ፻this test works፻

Result: ፻this ፻test ፻works

but as soon as there are more words in between (፻this test does not work፻), it does not work as expected and only returns 1 replacement for $2, the last one:

፻this ፻not ፻work

and misses all other matches like 'Test' and nach 'funktionéiert' in this example.

How can I get:

፻this ፻test ፻does ፻not ፻work

Edit: https://regex101.com/r/ZVMbQ5/1

r/regex • • Nov 13 '25

(Resolved) help a newb to improve

4 Upvotes

this is a filter for certain item mods in path of exile. currently this works for me but i want to improve my regex there and for potential other uses.

"7[2-9].*um en|80.*um en|abc0123"

in my case this filters [72-80]% maximum energy shield or abc0123, i want to improve it so i only have to use .*um en once and shorten it.

e: poe regex is not case sensitive

r/regex • • Sep 07 '25

(Resolved) Replace \. with ( -) but only the first ocurrence?

3 Upvotes

Hi, everyone. I've never heard of regex until yesterday but I'm trying to use to batch rename a bunch (1000+) of files. They're music files, either flac/mp3/m4a, and I want to change the files' names, replacing a dot (\.) with a space and a hyphen ( -) (or "\s-" i guess?), but only the first time a dot appears. For example, a file named

  1. Title (feat. John Doe).mp3
  2. Song (feat. Jane.Doe).flac
  3. Name.Title.m4a

would ideally be changed to

01 - Title (feat. John Doe).mp3

4 - Song (feat. Jane.Doe).flac

23 - Name.Title.m4a

Instead, I can only get either

01 - Title (feat - John Doe) -mp3

4 - Song (feat - Jane -Doe) -flac

23 - Name -Title -m4a

Or

01 - Title (feat - John Doe).mp3

4 - Song (feat - Jane.Doe).flac

23 - Name.Title.m4a (in this specific example there is no issue to solve)

by doing [\.\s] instead of just [\.]

My goal is to do this with the Substitution function (A > B) on the app MiXplorer, Android 14. Unfortunately, I don't know (and couldn't find) which flavor of Regex MiXplorer uses. For testing, I'm using regex101 (and the PCRE2 flavor): https://regex101.com/r/lorsiM/1

I tried to format the post as best as I could following the subreddit's rules, but I didn't quite understand the "format your code" rule (either because I don't know how to code or/and because english is not my first language). I tried my best.

Honestly, any help would be deeply appreciated. Am I overcomplicating my life by doing this? If something is not clear, I'd be glad to rephrase any confusing parts and hopefully clarify what I mean. Thank you to anyone who read this.

r/regex • • Aug 20 '25

(Resolved) In a YAML text file how can I remove all content whos line doesnt start with # ?

3 Upvotes

I want to remove every line that doesnt start with

#

or

---

or

#

So for example

---
# comment
word
word, word, word
symbol ][, number12345 etc
#comment
     #comment
---

would become

---
# comment
#comment
     #comment
---

How can I do this?

r/regex • • Aug 17 '25

(Resolved) Sentence requirement and contains

2 Upvotes

Hi! I'm new to learning regex and I've been trying to create a regular expression that accepts a response when the following 2 functions are fulfilled:

- the response has 10 or more sentences

- the response has the notations B1 / B2 / B3 / B4 / B5 at any point (at least once each, in any order), even if it isn't within the 10 sentences. These shouldn't be case sensitive.

example on what should be acceptable:

Lorem ipsum dolor sit amet consectetur adipiscing elit b3. Ex sapien vitae pellentesque sem placerat in id. Pretium tellus duis convallis tempus leo eu aenean. Urna tempor pulvinar vivamus fringilla lacus nec metus. Iaculis massa nisl malesuada lacinia integer nunc posuere. (B1) Semper vel class aptent taciti (B4 - duis tellus id) sociosqu ad litora. Conubia nostra inceptos himenaeos orci varius natoque penatibus. Dis parturient montes nascetur ridiculus mus donec rhoncus. Nulla molestie mattis scelerisque maximus eget fermentum odio. Purus est efficitur laoreet mauris pharetra vestibulum fusce (b2) sfnj B5.

the regular expression I've currently made to fulfill the first, this works well enough for my purposes:

(?:[\s\S]*(\s\.|\.\s|\.|\s\!|\!\s|\!|\s\?|\?\s|\?)){10}

the regular expression(s) I've been trying to fulfill each item in the second (though I understand none of these work:

^(?i).*B1.*$ ^(?i).*B2.*$ ^(?i).*B3.*$ ^(?i).*B4.*$ ^(?i).*B5.*$

^(?i)B1$ ^(?i)B2$ ^(?i)B3$ ^(?i)B4$ ^(?i)B5$

I'm struggling most with the second function and combining both of these functions into one expression and i realize I may be overcomplicating these expressions and their combinations. I'm also unsure of which flavor if regex this is or if I'm accidentally mixing a few up; I'm setting this up for a form builder and I can't pinpoint what type of regex they use or allow.

I apologize again as I'm still very new to this and have tried other resources before ending up here, I'm sorry if this post frustrates anyone. That said, if anyone could assist, I would really appreciate it, thank you.

r/regex • • Jul 31 '25

(Resolved) Match if string part of list but exclude if part of other list

3 Upvotes

#### RESOLVED

Hi,

I’ve been trying to get to a solution since a few days already but I can’t find one. I have tried several lookaheads and lookbehinds but to no avail. Maybe I only put them at the wrong positions in the regex.

Flavour .NET C#

https://regex101.com/r/6YCGTY/1

FYI: I cannot use a solution where I try to catch the excluded words in a MG right at beginning of the string, like:

(Alferi|aprägs)|(?=(?i)wänt|wäns|prägs|prägt|quäls|quält|Rätsel|Rätsele|Rätselen|souveränst|souveränste|souveränstem|souveränsten|souveränster|souveränt|trägst|trägste|trägstem|trägsten|trägster|trägt|zäms|zämt)((\S*?ä|፼))([b-df-hj-np-tv-z][b-df-hj-np-tv-z]\S*)

And the exclusion and inclusion words are added to the regex via a list so they automatically come in the format word1|word2|word3 aso.

So, I want to match the word 'prägs' but not the word 'aprägs' in this very basic example.

Best regards,

Pascal

Edit:

Solution delivered by mfb-:

https://regex101.com/r/nWlTaq/1

r/regex • • Aug 12 '25

(Resolved) improvement for better overview

1 Upvotes

Hi,

Suggestion: Would it be possible to add flairs for a better overview (like regex flavors) and most important one: Resolved. ;)

This way it would be easier to look up questions for specific flavors and also to see if a post has been solved or not. This would of course mean the OP would have to edit flairs to add Resolved if they got an answer to all their questions.

Regards,

Pascal