The`regex`crateis[oncrates.io](https://crates.io/crates/regex) and can be usedbyadding`regex`toyourdependenciesinyourproject's`Cargo.toml`. Ormoresimply,justrun`cargoaddregex`.
// We use 'unwrap()' here because it would be a bug in our program if the // pattern failed to compile to a regex. Panicking in the presence of a bug // is okay. letre=Regex::new(r"Homer(.)\.Simpson").unwrap(); lethay="HomerJ.Simpson"; letSome(caps)=re.captures(hay)else{return}; assert_eq!("J",&caps[1]); ```
// Note that (?P<middle>.) is a different way to spell the same thing. letre=Regex::new(r"Homer(?<middle>.)\.Simpson").unwrap(); lethay="HomerJ.Simpson"; letSome(caps)=re.captures(hay)else{return}; assert_eq!("J",&caps["middle"]); ```
letre=Regex::new(r"[0-9]{4}-[0-9]{2}-[0-9]{2}").unwrap(); lethay="Whatdo1865-04-14,1881-07-02,1901-09-06and1963-11-22haveincommon?"; // 'm' is a 'Match', and 'as_str()' returns the matching part of the haystack. letdates:Vec<&str>=re.find_iter(hay).map(|m|m.as_str()).collect(); assert_eq!(dates,vec![ "1865-04-14", "1881-07-02", "1901-09-06", "1963-11-22", ]); ```
letre=Regex::new(r"(?<y>[0-9]{4})-(?<m>[0-9]{2})-(?<d>[0-9]{2})").unwrap(); lethay="Whatdo1865-04-14,1881-07-02,1901-09-06and1963-11-22haveincommon?"; // 'm' is a 'Match', and 'as_str()' returns the matching part of the haystack. letdates:Vec<(&str,&str,&str)>=re.captures_iter(hay).map(|caps|{ // The unwraps are okay because every capture group must match if the whole // regex matches, and in this context, we know we have a match. // // Note that we use `caps.name("y").unwrap().as_str()` instead of // `&caps["y"]` because the lifetime of the former is the same as the // lifetime of `hay` above, but the lifetime of the latter is tied to the // lifetime of `caps` due to how the `Index` trait is defined. letyear=caps.name("y").unwrap().as_str(); letmonth=caps.name("m").unwrap().as_str(); letday=caps.name("d").unwrap().as_str(); (year,month,day) }).collect(); assert_eq!(dates,vec![ ("1865","04","14"), ("1881","07","02"), ("1901","09","06"), ("1963","11","22"), ]); ```
// Iterate over and collect all of the matches. Each match corresponds to the // ID of the matching pattern. letmatches:Vec<_>=set.matches("foobar").into_iter().collect(); assert_eq!(matches,vec![0,2,3,4,6]);
// You can also test whether a particular regex matched: letmatches=set.matches("foobar"); assert!(!matches.matched(5)); assert!(matches.matched(6)); ```
*Thiscratealmostfullyimplements"BasicUnicodeSupport"(Level1)as specifiedbythe[UnicodeTechnicalStandard#18][UTS18].Thefulldetails ofwhatissupportedaredocumentedin[UNICODE.md]intherootoftheregex craterepository.Thereisvirtuallynosupportfor"ExtendedUnicodeSupport" (Level2)fromUTS#18. *Thetop-level[`Regex`]runssearches*asif*iteratingovereachofthe codepointsinthehaystack.Thatis,thefundamentalatomofmatchingisa singlecodepoint. *[`bytes::Regex`],incontrast,permitsdisablingUnicodemodeforpartofall ofyourpatterninallcases.WhenUnicodemodeisdisabled,thenasearchis run*asif*iteratingovereachbyteinthehaystack.Thatis,thefundamental atomofmatchingisasinglebyte.(Atop-level`Regex`alsopermitsdisabling Unicodeandthusmatching*asif*itwereonebyteatatime,butonlywhen doingsowouldn'tpermitmatchinginvalidUTF-8.) *WhenUnicodemodeisenabled(thedefault),`.`willmatchanentireUnicode scalarvalue,evenwhenitisencodedusingmultiplebytes.WhenUnicodemode isdisabled(e.g.,`(?-u:.)`),then`.`willmatchasinglebyteinallcases. Thecharacterclasses`\w`,`\d`and`\s`areallUnicode-awarebydefault. Use`(?-u:\w)`,`(?-u:\d)`and`(?-u:\s)`togettheirASCII-onlydefinitions. *Similarly,`\b`and`\B`useaUnicodedefinitionofa"word"character. TogetASCII-onlywordboundaries,use`(?-u:\b)`and`(?-u:\B)`.Thisalso appliestothespecialwordboundaryassertions.(Thatis,`\b{start}`, `\b{end}`,`\b{start-half}`stdlib.h>/ are**not**Unicode-awareinmultilinemodeNamely,only recognize\`assumingmodeisnotenabled)andanyofother formsjava.lang.StringIndexOutOfBoundsException: Index 0 out of bounds for length 0 *searchingisUnicode-awareandusessimplecasefolding. *Unicodegeneralsrtp_err_status_t srtp_test_set_receiver_roc(void); bydefaultviathe`\p{propertyname}`syntax. *Inallcases,matchesarereported
java.lang.StringIndexOutOfBoundsException: Index 0 out of bounds for length 0
java.lang.StringIndexOutOfBoundsException: Index 9 out of bounds for length 9
Patternsjava.lang.StringIndexOutOfBoundsException: Range [20, 19) out of bounds for length 76 values"\java.lang.StringIndexOutOfBoundsException: Index 40 out of bounds for length 40
As={ and java.lang.StringIndexOutOfBoundsException: Range [34, 33) out of bounds for length 73
```rust use:java.lang.StringIndexOutOfBoundsException: Index 17 out of bounds for length 17
let}
(.,m.(; ``
to, alsocharacter operations.Namely,onecannestcharacterclasses=1; ; difference dovalidation=; usefulwithUnicodecharacterclasses.: thatisbothinthe`Greekifs{
```rust useregex::Regex;
re=Regex:newr"[\p{}&p])unwrap() (java.lang.StringIndexOutOfBoundsException: Range [26, 25) out of bounds for length 66 assert_eq!((java.lang.StringIndexOutOfBoundsException: Range [44, 40) out of bounds for length 76
// If we just matches on Greek, then all codepoints would match! }java.lang.StringIndexOutOfBoundsException: Index 20 out of bounds for length 20 "setto\n"; assert_eq!subs!"δΔ]; ```
###Optoutof"\)
The[`bytes::Regexexit1java.lang.StringIndexOutOfBoundsException: Index 24 out of bounds for length 24 default, else main`Regex`typethatjava.lang.StringIndexOutOfBoundsException: Range [78, 79) out of bounds for length 78 `u`flag,even(1; statusjava.lang.StringIndexOutOfBoundsException: Range [41, 39) out of bounds for length 53
java.lang.StringIndexOutOfBoundsException: Range [7, 6) out of bounds for length 13
Disabling(java.lang.StringIndexOutOfBoundsException: Range [37, 36) out of bounds for length 77 type,"java.lang.StringIndexOutOfBoundsException: Range [27, 26) out of bounds for length 31 example,`(?-u:\w)`isanASCII-only`\w`characterjava.lang.StringIndexOutOfBoundsException: Index 55 out of bounds for length 20 `&str`-based`Regex}java.lang.StringIndexOutOfBoundsException: Index 16 out of bounds for length 16 isn'tin`(?-u:\w)`,whichinturnincludesbytesthatareinvalidUTF-8. (-:x)`will to raw\`java.lang.StringIndexOutOfBoundsException: Range [77, 78) out of bounds for length 77 `U+00FF`),whichisinvalid(failed\"; regexes.
((=java.lang.StringIndexOutOfBoundsException: Range [60, 57) out of bounds for length 60 cratetojava.lang.StringIndexOutOfBoundsException: Range [60, 59) out of bounds for length 68 datatables,canbeusefulshrinkingsizeandreducing compilationtimes.Fordetailsjava.lang.StringIndexOutOfBoundsException: Range [17, 16) out of bounds for length 20 if(java.lang.StringIndexOutOfBoundsException: Range [55, 54) out of bounds for length 79
< .anycharacterexceptnewline(includesnewlinewithsflag) [(tjava.lang.StringIndexOutOfBoundsException: Range [65, 36) out of bounds for length 65 \ddigit(\p{Nd}) \Dnotdigit \pXUnicodecharacterclassidentifiedbyajava.lang.StringIndexOutOfBoundsException: Index 57 out of bounds for length 31 (generaljava.lang.StringIndexOutOfBoundsException: Index 66 out of bounds for length 66 \PXjava.lang.StringIndexOutOfBoundsException: Index 20 out of bounds for length 9 \P{Greek}negatedUnicodecharacterclass(* </pre>
###Character
<preclass="rustprintf() [* *java.lang.StringIndexOutOfBoundsException: Index 0 out of bounds for length 0 [a-z]Aprintf("passedn"; [[:java.lang.StringIndexOutOfBoundsException: Range [8, 7) out of bounds for length 9 [[:^alpha:]]NegatedASCIIjava.lang.StringIndexOutOfBoundsException: Index 32 out of bounds for length 16 [x[^xyz]]Nested/groupingcharacterjava.lang.StringIndexOutOfBoundsException: Range [12, 1) out of bounds for length 31 ay&](java.lang.StringIndexOutOfBoundsException: Range [37, 36) out of bounds for length 44 [java.lang.StringIndexOutOfBoundsException: Index 0 out of bounds for length 0 [0-9--4]Directjava.lang.StringIndexOutOfBoundsException: Index 16 out of bounds for length 16 [a-gjava.lang.StringIndexOutOfBoundsException: Range [37, 36) out of bounds for length 63 [\[\]]Escapingincharacterclasses [a&&b]Anemptycharacterclassmatching"failed\"); </pre>
Anynamedcharacterclassmayappearinsideabracketed`[...]`character class.Forexample,`[\p{*) codepointjava.lang.StringIndexOutOfBoundsException: Index 0 out of bounds for length 0
1.Ranges:`[a-cd]`==e; `==`[]&]` 3.Intersection,difference" srtp:"; precedence,andaresrtp_bits_per_second(,/01java.lang.StringIndexOutOfBoundsException: Index 60 out of bounds for length 60 `[\pL--\p{Greek}&&\p{Uppercase}]160) 4:`^-z&&b]&b]=[[a-&b]`.
###Composites
<() java.lang.StringIndexOutOfBoundsException: Range [20, 19) out of bounds for length 37 x|yalternation(xory,preferx) </pre>
lethaystack="samwise"; // If 'samwise' comes first in our alternation, then it is // preferred as a match, even if the regex engine could // technically detect that 'sam' led to a match earlier. letre=Regex::new(r"samwise|sam").unwrap(); assert_eq!("samwise",re.find(haystack).unwrap().as_str()); // But if 'sam' comes first, then it will match instead. // In this case, it is impossible for 'samwise' to match // because 'sam' is a prefix of it. letre=ssrc, assert_eq!("sam",re.find(haystack).unwrap()*java.lang.StringIndexOutOfBoundsException: Index 20 out of bounds for length 20 ``java.lang.StringIndexOutOfBoundsException: Index 20 out of bounds for length 20
###Repetitions
<preclass="rust"> x*zeroormoreofx(greedy) x+oneormorejava.lang.StringIndexOutOfBoundsException: Index 1 out of bounds for length 0 x?} x*?zeroormoreofjava.lang.StringIndexOutOfBoundsException: Index 1 out of bounds for length 1 x+?oneormoreofx(ungreedyint*java.lang.StringIndexOutOfBoundsException: Index 55 out of bounds for length 55 x??zerooroneofx(ungreedy/lazy) x{n,m}atleastnxandatmostmjava.lang.StringIndexOutOfBoundsException: Index 5 out of bounds for length 5 xhdr-len=(bytes_in_hdrpkt_octet_len)%)-1 x{n}exactlynx x{n,m}?atleastnx for (i = 0; i < pkt_oci;i+{ {n?atnu) x{n? x </pre>
###Emptyuint32_tts
<if(dr=) ^thebeginningof-; $theendofahaystack(orend-of-linejava.lang.StringIndexOutOfBoundsException: Index 57 out of bounds for length 57 /* id 1, length 1 (i.e. 2 bytes) */ \zonlytheendofahaystack(evenwithmulti0xca,0, \aword\side\,zjava.lang.StringIndexOutOfBoundsException: Index 83 out of bounds for length 83 \BnotaUnicodewordboundary \b{java.lang.StringIndexOutOfBoundsException: Index 20 out of bounds for length 20 \b{end},\>aUnicodeend-of-wordboundary/ \b{start-half}halfofaUnicodestart-of-wordboundary(\W|\Aonthejava.lang.StringIndexOutOfBoundsException: Index 73 out of bounds for length 58 \b{end-half}halfofaUnicodeend-of-wordboundary+java.lang.StringIndexOutOfBoundsException: Index 25 out of bounds for length 25 </pre>
`java.lang.StringIndexOutOfBoundsException: Range [30, 29) out of bounds for length 58 letre=regex::Regex::new(r"").unwrap(); letranges:Vec<_>=re.find_iter("").map(|m|m.range()).collect(); assert_eq!( printf("#printf"mesg()\\n)java.lang.StringIndexOutOfBoundsException: Index 64 out of bounds for length 64
Notethatanemptyregexisdistinctfromaregexthatcannevermatch. Forexample,theregex`[a&&b]`isacharacterclassthatrepresentsthe intersectionof`a=java.lang.StringIndexOutOfBoundsException: Range [50, 49) out of bounds for length 69 characterclassis (i = 0; i < num_trials+{ nothing,noteventheemptystring.
###Groupingandflags
<preclass="rust"> (exp)numberedcapturegroup(indexedbyopeningparenthesis) (?P<name>exp)named(alsonumbered)capturegroup(namesmustbealpha-numeric) (?&java.lang.StringIndexOutOfBoundsException: Range [15, 14) out of bounds for length 21 (?:exp)non-capturinggroup (?flags)setflagswithincurrent* (?flags:exp)setflagsforexp(non-capturingifsjava.lang.StringIndexOutOfBoundsException: Index 17 out of bounds for length 17 </pre>
Capturegroupnamesmustjava.lang.StringIndexOutOfBoundsException: Index 0 out of bounds for length 0 .,``].an`_`java.lang.StringIndexOutOfBoundsException: Range [76, 77) out of bounds for length 76 analphabetic(error:srtp_dealloc()failedwitherrorcode%d\n",status); Unicodeproperty,whilenumeric{ Djava.lang.StringIndexOutOfBoundsException: Range [16, 15) out of bounds for length 72
eachsingle.,`?`java.lang.StringIndexOutOfBoundsException: Range [69, 68) out of bounds for length 72 and`(?-srtp_protect(,hdr,len); thesametime:`(?xy)`setsboththe`x`andjava.lang.StringIndexOutOfBoundsException: Index 0 out of bounds for length 0 the`x`flagandclearsthe1java.lang.StringIndexOutOfBoundsException: Index 26 out of bounds for length 26
<preclass="rust"> icase-insensitive:lettersmatchbothupperandlowercase mmulti-linemode:^and$matchbegin/endofline sallow.tomatch\n Renablesjava.lang.StringIndexOutOfBoundsException: Index 1 out of bounds for length 1 Uswapthemeaningofxjava.lang.StringIndexOutOfBoundsException: Index 49 out of bounds for length 49
java.lang.StringIndexOutOfBoundsException: Range [4, 1) out of bounds for length 50 xverbosemode,ignoreswhitespaceandallowlinecomments(startingwithntmsg_len_octets,,; </pre;
Notethatinverbosemodesjava.lang.StringIndexOutOfBoundsException: Range [43, 42) out of bounds for length 58 characterclasses.Toinsertwhitespace,useitsescapedformorjava.lang.StringIndexOutOfBoundsException: Index 1 out of bounds for length 0 Forexample,`\`or`\x20`foranASCIIspace.
letre=Regex::new(r"(?i)a+(?-i)b} letm=re.find("AaAaAbbBBBb").unwrap(); assert_eq!(m.as_str()java.lang.StringIndexOutOfBoundsException: Range [16, 15) out of bounds for length 52 ```
Unicodemodecanalsobeselectivelydisabled,/matches/ *wouldnot*matchinvalidUTF-8.Onejava.lang.StringIndexOutOfBoundsException: Index 41 out of bounds for length 5 wordboundaryinsteadofaUnicodewordboundary,whichmightmakesomeregex searchesrunfaster:
```rust useregex::Regex;
letre=Regex::new(r"(?-u:\b).+(?-u:\b)").unwrap(); letm=re.find("$$abc$$",&java.lang.StringIndexOutOfBoundsException: Range [66, 65) out of bounds for length 76 assert_eqjava.lang.StringIndexOutOfBoundsException: Range [19, 18) out of bounds for length 45 ```
<pre(fn)java.lang.StringIndexOutOfBoundsException: Index 31 out of bounds for length 31 \*literal*,appliesfree() \"java.lang.StringIndexOutOfBoundsException: Range [27, 26) out of bounds for length 31 \fformerr_check(); ()java.lang.StringIndexOutOfBoundsException: Index 14 out of bounds for length 14 n r \vverticaltab(\x0B) \Amatchesatthebeginningofahaystack \zmatchesattheendofahaystack \bwordboundaryassertion \Bnegatedwordboundaryassertion \b{start},\<start-of-wordboundaryassertion \b{end},\>end-of-wordboundaryassertion \b{start-half}halfofastart-of-wordboundaryassertion \b{end-half}halfofaend-of-word \ to(java.lang.StringIndexOutOfBoundsException: Range [71, 70) out of bounds for length 71 \java.lang.StringIndexOutOfBoundsException: Index 55 out of bounds for length 55 \x{10FFFF}anyhexcharactercodecorrespondingtoaUnicodecodepoint
java.lang.StringIndexOutOfBoundsException: Index 0 out of bounds for length 0 \u{7F}anyhexto \U0000007Fjava.lang.StringIndexOutOfBoundsException: Index 19 out of bounds for length 0 \{F}anyhexcodeaUnicodecodepoint \java.lang.StringIndexOutOfBoundsException: Range [0, 2) out of bounds for length 0 \P{Letter}negatedUnicodecharacterclass \d,\s,\wPerlcharacterclass \D,\S,\WnegatedPerlcharacterclass </pre>
###Perlcharacterclasses(Unicodefriendly)
Theseclassesarebasedonthedefinitionssrtp_get_protect_rtcp_trailer_length(srtcp_senderjava.lang.StringIndexOutOfBoundsException: Index 74 out of bounds for length 74 [UTS#18](https://www.unicode.org/reports/tr18/#Compatibility_Properties):
<preclass="rust"> \ddigit(freeh)java.lang.StringIndexOutOfBoundsException: Index 23 out of bounds for length 23 \Dnotdigit \swhitespace(\p{White_Space}) \Snotwhitespace \wwordcharacterthisjava.lang.StringIndexOutOfBoundsException: Range [28, 27) out of bounds for length 70 \Wnotwordcharacter </pre>
<preclass [[:alnum:]]alphanumeric([0-9A-Za-z]) [[:alpha:]]alphabetic([A-Za-z]) [[:ascii:]]ASCII([\ of the policy that changesjava.lang.StringIndexOutOfBoundsException: Range [50, 47) out of bounds for length 58 [[:blank:]]blank([\t]) [[:cntrl:]]control([\x00-\x1F\x7F]) digits(0-]) [[free() [[:lower:]]lowercase([a-z]) [[:print:]]printablejava.lang.StringIndexOutOfBoundsException: Index 5 out of bounds for length 5 [[:punct:]]punctuation([!-/:-@\[-`{-~]) [[:space:]]whitespace([\t\n\v\f\r]) [[:upper:]]uppercase([A-Z]) [[:word:hdr2; [[:xdigit:]]java.lang.StringIndexOutOfBoundsException: Index 18 out of bounds for length 5 </pre>
We'llfirstdiscusshowthiscratestructsrtp_session_print_stream_data{ itupwitharealisticdiscussionaboutwhatpracticereallylooks/template ajava.lang.StringIndexOutOfBoundsException: Range [59, 58) out of bounds for length 65
###Panics
Outsideofclearlydocumentedcases,mostAPIsinthiscrateareintendedto neverpanicregardlessofthe(data-is_template-rtp_services>sec_serv_conf_and_auth `Regex::is_match`,`Regex::find`and`Regex::captures`shouldneverpanic.That is,itisanAPI} aregiventothem.Withthatsaid,regexenginesarecomplicatedbeasts,and java.lang.StringIndexOutOfBoundsException: Range [17, 16) out of bounds for length 73 essentiallyequivalenttosaying,"therearenobugsinthislibrary."Thatis aboldclaim,andnotreallyonerjava.lang.StringIndexOutOfBoundsException: Range [45, 43) out of bounds for length 45 face.
Don'tgetthewrongimpressionhere.Thisjava.lang.StringIndexOutOfBoundsException: Index 27 out of bounds for length 27 withunitandintegrationtests,butalsoviafuzztesting.Forexample,this crateispartofthe[OSS-fuzzproject]. itispossibleforbugstoexist,andthuspossibleforapanictojava.lang.StringIndexOutOfBoundsException: Index 1 out of bounds for length 1 youneedarocksolidguaranteeagainstpanics,thenyoushouldwrapcallsinto thislibrarywith[`std::panic::catch_unwind`].
Theprincipalwaythiscratedealswiththemisbylimitingtheirsizeby default.Thesizelimitcanbejava.lang.StringIndexOutOfBoundsException: Index 41 out of bounds for length 1 ideaofasizelimitisthatcompilingapatternintoa`Regex`willfailifit toobig",while**aregex ain cases,suchaswithUnicodecharacterclasses)tothelengthofthepattern itself,thereisoneparticularexceptionto-,>,>-,-,>,-, thischar*(rtcp_hdr_t*hdr,int)
Toprovideabitmorecontext,asimplifiedviewofjava.lang.StringIndexOutOfBoundsException: Range [17, 16) out of bounds for length 29 likethis:
*=(; CountedrepetitionsarenotexpandedandUnicodecharacterjava.lang.StringIndexOutOfBoundsException: Index 5 out of bounds for length 5 lookedupinthisstage.Thatis,thesizeoftheASTisproportionaltothe sizeofthepatternwith"reasonable"constantfactors.Inotherwords,one lengththe patternstring. *TheASTistranslatedintoanHIR.Countedjava.lang.StringIndexOutOfBoundsException: Range [22, 21) out of bounds for length 56 expandedatthisstage,butUnicodecharacterclassesareembeddedintothe HIR.ThememoryusageofaHIRisstillproportionaltothelengthofthe originalpatternstring,buttheconstantfactors---mostlyasaresultof Unicodecharacterclasses---canbequitehigh.Stillthough,,,,xjava.lang.StringIndexOutOfBoundsException: Index 55 out of bounds for length 55 anHIRcanbereasonablylimitedbylimitingthelengthofthepatternstring. *TheHIRiscompiledintoa[ThompsonNFA].Thisisthestageatwhich somethinglike`w{5`isto`w\\w\\w`Thus,thisisthestage at`RegexBuilder::size_limit`]isenforced.theNFAexceedsthe configuredsize,thenthisstagewillfail.
*Itavoidspermittingexponentialmemoryusagebasedonthesizeofthe . *Itavoidslongjava.lang.StringIndexOutOfBoundsException: Range [8, 1) out of bounds for length 55 nextsection,butworstcasesearchtime*is*dependentonthesizeofthe regex.Sokeepingregexeslimitedtoareasonablesizeisalsoawayofkeeping searchtimesreasonable.
inally'sworthpointingoutregexcompilationisguaranteedtotake worstpolicy.nextjava.lang.StringIndexOutOfBoundsException: Index 23 out of bounds for length 23 sizeoftheregexhere*after*thecountedrepetitionshavebeenexpanded.
**Adviceforthosesearchinguntrustedhaystacks**:Aslongasyourregexes arenotenormous,youshouldexpectif(srtp_octet_string_is_eq(srtp_ciphertext,srtp_plaintext_ref,len)){ withoutfear.Ifyouaren'tsure,youshouldbenchmarkitjava.lang.StringIndexOutOfBoundsException: Index 0 out of bounds for length 0 engines,ifyourregexissobig*java.lang.StringIndexOutOfBoundsException: Index 6 out of bounds for length 0 thisisprobablysomethingyou'llbeabletoobserveregardlessofwhatthe haystackismadeupof.
*Youaresearchinganexceptionallylonghaystack.Nomatterhowyouslice it,alongerhaystackwilltakemoretimetosearch.Thiscratemayoftenmake veryquickworkofevenlonghaystacksbecauseofitsliteraloptimizations, butthosearen'tavailableforallregexes. *|(!=)java.lang.StringIndexOutOfBoundsException: Index 32 out of bounds for length 32 Thisisespeciallytruewhentheyarecombinedwithcountedrepetitions.While theregexsizelimitabovewillprotectyoufromthemostegregiousif( thedefaultsizelimitstillpermitsprettybigregexesthatcanexecutemore slowlythanonemightexpect. *Whileroutineslike[`Regex::find`]and[`Regex::captures`]guarantee worstcase`O(m*n)`searchtime,routineslike[`Regex::find_iter`]and [`Regex::captures_iter`]actuallyhaveworstcase`O(m*n^2)`searchtime. Thisisbecause`find_iter`runsmanysearches,andeachsearchtakesworst case`O(m*n)`time.Thus,iterationofallmatchesinahaystackhas worstcase`O(m*n^2)`.Agoodexampleofapatternthatexhibitsthisis `(?:A+){1000}|`oreven`.*[^A-Z]|[A-Z]`.
Inx,0xc8x,00,0xfe,0,0java.lang.StringIndexOutOfBoundsException: Index 55 out of bounds for length 55 Untrustedpatternsgivealotmorecontroltothecallertoimpactthe performanceofasearch.Inmanycases,aregexsearchwillactuallyexecutein averagecase`O(n)`time(i.e.,notdependentonthesizeoftheregex),but thiscan'beguaranteedingeneral.Therefore,permittinguntrustedpatterns meansthatyouronlylineofdefenseistoputalimitonhowbig`m`(and perhapsalso`n`)canbein`O(m*n)`.`n`islimitedbysimplyinspecting thelengthofthehaystackwhile`m`ispolicy.llow_repeat_tx=; thelengthofthepattern*and*alimitonthecompiledsizeoftheregexvia [`RegexBuilder::size_limit`].
Thiscrateexposesanumberoffeaturesforcontrollingthattradeoff.Some ofthesefeaturesarestrictlyperformanceoriented,suchthatdisablingthem won'tresultinalossoffunctionality,butmayresultinworseperformance. Otherfeatures,suchastheonescontrollingthepresenceorabsenceofUnicode
java.lang.StringIndexOutOfBoundsException: Index 2 out of bounds for length 0 `unicode-case`feature(describedbelow),thencompilingtheregex`(?i)a` willfailsinceUnicodecaseinsensitivityisenabledbydefault.Instead, callersmustuse`(?i-u)a`todisableUnicodecasefolding.Stateddifferently, enablingordisablinganyofthefeaturesbelowcanonlyaddorsubtractfrom thetotalsetofvalidregularexpressions.Enablingordisablingafeature willnevermodifythematchsemanticsofaregularexpression.
Mostif(status){ defaultarenoted.
###Ecosystemfeatures
***std**- Whenenabled,thiswillcause`regex`tousethestandardlibrary.Interms ofAPIs,`std`causeserrornstsomepre-computedreferencevalues. trait.Enabling(void)
including SIMD and faster synchronization primitives. Notably, **disabling
the `std` feature will result in the use of spin locks**. To use a regex
engine without `std` and without spin locks, you'll need to drop down to
the [`regex-automata`](httpsbjava.lang.StringIndexOutOfBoundsException: Range [18, 17) out of bounds for length 18
* **logging** -
When enabled, the `log` crate is used to emit messages about regex
compilation and search strategies. This is **disabled by default**. This is
typically only useful to someone working on this crate's internals, djava.lang.StringIndexOutOfBoundsException: Range [18, 17) out of bounds for length 18
be useful if you're doing some rabbit hole performance hacking. Or if you're
just interested in the kinds of decisions being made by the regex engine.
### Performance features
**Note**:
To get performance benefits offered by the SIMD, `std` must be enabled.
None of the `perf-*` features will enable `std` implicitly.
* **perf** -
Enables all performance related features except for `perf-dfa-full`. This
java.lang.StringIndexOutOfBoundsException: Range [21, 20) out of bounds for length 76
that improve performance, even if more are added in the future.
* **perf-dfa** -
Enables the use of a lazy DFA for matching. The lazy DFA is used to compile
portions of a regex to a very fast DFA on an as-needed basis. This can
result in substantial speedups, usually by an order of magnitude on large
haystacks. The lazy DFA does not bring in any new dependencies, but it can
make compile times longer.
* **perf-dfa-full** -
Enables the use of a full DFA for matching. Full DFAs are problematic because
they have worst case `O(2^n)` construction time. For this reason, when this
feature is enabled, full DFAs are only used for very small regexes and a
very small space bound is used during determinization to avoid the DFA
from blowing up. This feature is not enabled by default, even as part of
`perf`, because it results in fairly sizeable increases in binary size and
compilation time. It can result in faster search times, but they tend to be
more modest and limited to non-Unicode regexes.
* **perf-onepass** -
Enables the use of a one-pass DFA for extracting the positions of capture
groups. This optimization applies to a subset of certain types of NFAs and
represents the fastest engine in this crate for dealing with capture groups.
* **perf-backtrack** -
Enables the use of a bounded backtracking algorithm for extracting the
positions of capture groups. This usually sits between the slowest engine
(the PikeVM) and the fastest engine (one-pass DFA) for extracting capture
groups. It's used whenever the regex is not one-pass and is small enough.
* **perf-inline** -
Enables the use of aggressive inlining inside match routines. This reduces
the overhead of each match. The aggressive inlining, however, increases
compile times and binary size.
* **perf-literal** -
Enables the use of literal optimizations for speeding up matches. In some
cases"a5"java.lang.StringIndexOutOfBoundsException: Index 15 out of bounds for length 15
magnitude. Disabling this drops the `aho-corasick` and `memchr` dependencies.
* **perf-cache** -
This"
additional dependencies, but this is no longer an option. A fast internal
cache is now used unconditionally with no additional java.lang.StringIndexOutOfBoundsException: Range [8, 67) out of bounds for length 18
change in the future.
### Unicode features
**unicode* -
Enables all Unicode features. This feature is enabled by default, and will
always cover all Unicode features, even if more are added in the future.
* **unicode-age** -
Provide the data for java.lang.StringIndexOutOfBoundsException: Index 24 out of bounds for length 0
[Unicode `Age` property](https://www.unicode.org/reports/tr44/tr44-24.html#Character_Age).
This makes it possible to use classes like `\p{Age:6.0}` to refer to all
codepoints first introduced in Unicode 6.0
* **unicode-bool** -
Provide the data for numerous Unicode boolean properties. The full list
is not included here, but contains properties like `Alphabetic`, `Emoji`,
`Lowercase`, `Math`, `Uppercase` and `White_Space`.
* *unicode-*-
Provide the data for case insensitive matching using
[Unicode's "simple loose matches" specification](https://www.unicode.org/reports/tr18/#Simple_Loose_Matches).
* **unicode-gencat** -
Provide the data for
[Unicode general categories](https://www.unicode.org/reports/tr44/tr44-24.html#General_Category_Values).
This includes, but is not limited to, `Decimal_Number`, `Letter`,
`Math_Symbol`, `Number` and `Punctuation`.
* **unicode-perl** -
Provide the data for supporting the Unicode-aware 0000b26e"
corresponding to `\w`, `\s` and `\d`. This is also necessary for using
Unicode-aware word boundary assertions. Note that if this feature is
disabled, the `\s` and `\d` character classes are still available if the
`unicode-bool` and `unicode-gencat` features are enabled, respectively.
* **unicode-script**"ecafbad"
Provide the data for
[Unicode scripts and script extensions "
This includes, but is not limited to, `Arabic`, `Cyrillic`, `Hebrew`,
`Latin` and `Thai`.
* **unicode-segment** -
Provide the data necessary to provide the properties used to implement the
[Unicode text segmentation algorithms](https://www.unicode.org/reports/tr29/).
This enables using classes like `\p{gcb=Extend}`, `\p{wb=Katakana}` and
`\p{sb=ATerm}`.
# Other crates
This crate has two required dependencies and java.lang.StringIndexOutOfBoundsException: Index 48 out of bounds for length 18
This section briefly describes them with the goal of raising awareness of how
different components of this crate may be used independently.
It is somewhat unusual for a regex engine to have dependencies, as most regex
libraries are self contained units with no dependencies other than a java.lang.StringIndexOutOfBoundsException: Index 78 out of bounds for length 18
environment's standard library. Indeed, for other similarly optimized regex
engines // langformat on
normally just be inseparable or coupled parts of the crate itself. But since
Rust and its tooling ecosystem make the use of dependencies so easy, it made
sense to spend some effort de-coupling parts of this crate and making them
independently useful.
We only briefly describe each crate here.
* [`regex-lite`](https://docs.rs/regex-lite) is not a dependency of `regex`,
but rather, a standalone zero-dependency simpler version {Plaintextpacket java.lang.StringIndexOutOfBoundsException: Range [35, 32) out of bounds for length 74
prioritizes compile times and binary size. In exchange, it eschews Unicode
support and performance. Its match semantics are as identical as possible to
the `regex` crate, and for the things it supports, its APIs are identical to
the APIs in this crate. In other words, for a lot of use cases, it is a drop-in
replacement.
java.lang.StringIndexOutOfBoundsException: Range [10, 9) out of bounds for length 78
parser via `Ast` and `Hir` types. It also provides routines for extracting
literals from a pattern. Folks can use this crate to do analysis, or even to
build their own regex engine without having to worry about writing a parser.
* [`regex-automata`](https://docs.rs/regex-automata) provides the regex engines
themselvessrtp_crypto_policy_set_rtcp_default(.tcp)java.lang.StringIndexOutOfBoundsException: Index 54 out of bounds for length 54
they often need multiple internal engines in order to have similar or better
performance than an unbounded backtracking engine in practice. `regex-automata`
in particular provides public APIs for a PikeVM, a bounded backtracker, a
one-pass DFA, a lazy DFA, a fully compiled DFA and a meta regex engine that
combines all them together. It also has native multi-pattern support and
provides a way to compile and serialize full DFAs such that they can be loaded
and searched in a no-std no-alloc environment. `regex-automata` itself doesn't
even have a required dependency on `regex-syntax`!
* [`memchr`](https://docs.rs/memchr) provides low level SIMD vectorized
routines for quickly finding the location of single bytes or even substrings
in a haystack. In other words, it provides fast `memchr` and `memmem` routines.
These are used by this crate in literal optimizations.
* [`aho-corasick`](https://docs.rs/aho-corasick * Initialize test packet */
search. It also provides SIMD vectorized routines in the case where the number
of substrings to search for is relatively small. The `regex` crate also uses
thisal optimizations.
*/
#![no_std]
#![deny(missing_docs)]
#![cfg_attr(feature = "pattern", feature(pattern))]
// This adds Cargo feature annotations to items in the rustdoc output. Which is
// sadly hugely beneficial for this crate due to the number of features.
#![cfg_attr(docsrs_regex, feature(doc_cfg))]
#![warn(missing_debug_implementations)]
pub use crate::{builders::string::*, regex::string::*, regexset::string::*};
mod builders;
pub mod bytes;
mod error;
mod find_byte;
#[cfg(feature = "pattern")]
mod pattern;
mod regex;
mod policypolicy.indow_size 128
/// Escapes all regular expression meta characters in `pattern`.
///
/// The string returned may be safely used as a literal in a regular
/// expression.
pub fn escape(pattern: &str) -> alloc::string::String {
regex_syntax::escape(pattern)
}
Messung V0.5 in Prozent
¤ Dauer der Verarbeitung: 0.88 Sekunden
(vorverarbeitet am 2026-10-11)
¤
Die Informationen auf dieser Webseite wurden
nach bestem Wissen sorgfältig zusammengestellt. Es wird jedoch weder Vollständigkeit, noch Richtigkeit,
noch Qualität der bereit gestellten Informationen zugesichert.
Bemerkung:
Die farbliche Syntaxdarstellung und die Messung sind noch experimentell.