core/oga - oga

Commit Graph

Author	SHA1	Message	Date
Yorick Peterse	45d84d31da	Renamed rspec helper files	2015-03-22 22:50:03 +01:00
Yorick Peterse	70e4942d3e	CSS parser spec for "+ b"	2015-03-21 01:23:00 +01:00
Yorick Peterse	6039e1dbeb	XPath parsing spec for axes with predicates	2015-03-21 01:23:00 +01:00
Yorick Peterse	62fa2a9cc5	Spec for XPath functions inside predicates.	2015-03-21 01:23:00 +01:00
Yorick Peterse	194d981996	XPath specs for paths with multiple members.	2015-03-21 01:22:59 +01:00
Yorick Peterse	fdcd712ffe	Don't use Array#uniq in NodeSet#initialize. Removing this makes the process of parsing larger XML documents a bit faster. The downside is that NodeSet#initialize will no longer filter out duplicate nodes, though this is not something Oga itself relies upon. Methods such as NodeSet#push still do ignore elements already present.	2015-03-21 01:22:59 +01:00
Yorick Peterse	f83c03aaec	Fixed typo in NodeSet spec.	2015-03-21 01:22:59 +01:00
Yorick Peterse	006ef4d51a	Port over most of the old XML error handling. Some messages are a bit different due to ruby-ll's error handling, other than that it's largely the same stuff as before.	2015-03-21 01:22:59 +01:00
Yorick Peterse	d8b9725b82	Fixed SAX parsing of XML attributes. This was utterly broken, mainly due to me overlooking it. There are now 2 new callbacks to handle this properly: * on_attribute: to handle a single attribute/value pair * on_attributes: to handle a collection of attributes (as returned by on_attribute) By default on_attribut returns a Hash, on_attributes in turn merges all attribute hashes into a single one. This ensures that on_element _actually_ receives the attributes as a Hash, instead of an Array with random nil/XML::Attribute values.	2015-03-21 01:22:59 +01:00
Yorick Peterse	605d565104	Use sax_parse_html for HTML documents. I suspect the only reason this test ever passed due to Racc's error handling. Either way this was using the wrong method.	2015-03-21 01:22:59 +01:00
Yorick Peterse	2ec91f130f	Lazy decoding of XML/HTML entities. Instead of decoding entities in the lexer we'll do this whenever XML::Text#text is called. This removes the overhead from the parsing phase and ensures the process is only triggered when actually needed. Note that calling #to_xml and/or the #inspect methods on a Text (or parent) instance will also trigger the entity conversion process. The new entity decoding API supports both regular entities (e.g. &) as well as codepoint based entities (both regular and hexadecimal codepoints). To allow safe read-only access to Text instances from multiple threads a mutex is used. This mutex ensures that only 1 thread can trigger the conversion process. Fixes #68	2015-03-05 23:00:43 +01:00
Yorick Peterse	3b2055a30b	Refactored handling of literal HTML elements. This ensures newlines can appear in <style> / <script> tags when using IOs as input.	2015-03-04 11:44:31 +01:00
Yorick Peterse	78e40b55c0	Handle parsing of HTML <style> tags. This basically re-applies the technique used for HTML <script> tags. With this extra addition I decided to rename/normalize a few things so it's easier to add any extra tags in the future. One downside of this setup is that the following will not be parsed by Oga: <style> </script> </style> The same applies to script tags containing a literal </style> tag. Since this particular case is rather unlikely to occur I'm OK with not supporting it as it _does_ simplify the lexer quite a bit. Fixes #80	2015-03-03 16:28:05 +01:00
Yorick Peterse	142b467277	Set parent of nodes set using Element#inner_text= This ensures that any text nodes created using Element#inner_text= have their parent node set correctly.	2015-03-03 13:13:05 +01:00
Yorick Peterse	874d7124af	Don't convert <script> text to XML entities. Fixes #79.	2015-03-02 17:32:19 +01:00
Yorick Peterse	9a586363e9	Added XML::Document#html?	2015-03-02 16:39:40 +01:00
Yorick Peterse	351b5ac004	Added spec for lexing inline HTML script tags. Related issue: #70	2015-03-02 16:20:06 +01:00
Yorick Peterse	47a3c5e7f8	Use describe/it instead of context/example. This keeps things consistent with the general testing guidelines in the Ruby community. This in turn should hopefully make my life easier as I don't have to tell people to use this rather odd stlye I was using before.	2015-01-08 23:01:53 +01:00
Yorick Peterse	746c8052dd	Remove all nodes when calling Element#inner_text= This fixes #64.	2014-12-14 23:32:43 +01:00
Dmitry Krasnoukhov	26baf89440	Add missing entities to the decode/encode lists	2014-11-21 01:53:11 +02:00
Yorick Peterse	cbb2815146	Support for inline doctype rules plus newlines. This adds support for lexing/parsing XML documents that use an IO as input _and_ contain doctype rules with newlines in them. This fixes #63.	2014-11-18 20:02:55 +01:00
Yorick Peterse	ad4f650c5d	Fixed XML entity encoding/decoding ordering. Thanks to @krasnoukhov for providing the initial patch, which this commit is largely based on. This fixes #49.	2014-11-17 22:39:43 +01:00
Yorick Peterse	cd86d5d294	Allow removal of element attributes.	2014-11-17 09:00:40 +01:00
Yorick Peterse	804646cc5e	Don't modify raw namespaces. When calling Element#available_namespaces the list of namespaces returned by Element#namespaces must not be modified.	2014-11-17 00:01:16 +01:00
Yorick Peterse	57adabc068	Ensure SAX after_element receives meaningful args This changes the behaviour of after_element when parsing documents using the SAX parsing API. Previously it would always receive a nil argument, which is kinda pointless. This commit changes that by making sure it receives a namespace name (if any) and the element name. This fixes #54.	2014-11-16 23:32:32 +01:00
Yorick Peterse	448ff56e38	Fixed CSS eval specs for nth-(first\|last)-of-type.	2014-11-15 18:27:26 +01:00
Yorick Peterse	b464815577	Fixed AST generation for nth-(first\|last)-of-type.	2014-11-15 18:27:15 +01:00
Yorick Peterse	9eead81a7c	Fixed AST for :only-of-type	2014-11-15 18:08:26 +01:00
Yorick Peterse	1c301d40e2	Properly fixed AST for first-of-type/last-of-type This requires keeping track of the current element being processed. This in turn allows the usage of count() + preceding-sibling/following-sibling.	2014-11-15 17:58:56 +01:00
Yorick Peterse	e559b4b89b	Corrected :first-of-type eval spec.	2014-11-15 17:20:17 +01:00
Yorick Peterse	f1d574f342	Evaluate XPath predicates for every context node. Instead of evaluating a predicate once for all context nodes, they should instead be evaluated separately per context node.	2014-11-15 00:31:44 +01:00
Yorick Peterse	6daa3e7a00	Reverted AST changes for first-of-type Functions can't be used in combination with axes, so I'll just need to fix the position() function to work properly.	2014-11-14 23:51:46 +01:00
Yorick Peterse	2d6a2be2e8	Revert "Fixed XPath AST for :last-of-type" Axes can't be used in combination with functions. This reverts commit `b0b572a584`.	2014-11-14 23:49:49 +01:00
Yorick Peterse	8f3553f8f1	Fixed eval specs of :first-of-type & :last-of-type	2014-11-14 23:27:52 +01:00
Yorick Peterse	b0b572a584	Fixed XPath AST for :last-of-type This should count following nodes, not merely the position.	2014-11-14 23:27:15 +01:00
Yorick Peterse	0128dc50ae	Fixed CSS evaluation of :first-of-type The old XPath "position() = 1" would work in Nokogiri due to the way they retrieve descendants. In Oga however this would simply always return the first node. To fix this Oga now counts the amount of preceding siblings that match the same full name.	2014-11-14 01:25:03 +01:00
Yorick Peterse	e3a26c5d15	Allow querying of nodes using CSS.	2014-11-14 01:05:29 +01:00
Yorick Peterse	d47ca19ffa	Remaining CSS evaluation specs.	2014-11-14 00:24:54 +01:00
Yorick Peterse	518bedc3a1	CSS evaluator specs for :nth-of-type	2014-11-14 00:16:35 +01:00
Yorick Peterse	abfe6e3d61	CSS evaluator specs for nth-last-of-type	2014-11-14 00:13:11 +01:00
Yorick Peterse	c874ceabb9	CSS evaluator specs for :nth-last-child	2014-11-13 23:54:47 +01:00
Yorick Peterse	5964a2cda4	CSS evaluator specs for :nth-child(n)	2014-11-13 22:50:04 +01:00
Yorick Peterse	3237617bf5	CSS eval specs for various pseudo classes. This includes the following pseudos: * :empty * :first-child * :first-of-type * :last-child * :last-of-type	2014-11-13 10:10:31 +01:00
Yorick Peterse	27ffa4d3d5	Added various failing following/preceding specs.	2014-11-13 01:11:13 +01:00
Yorick Peterse	97a9a11db1	Failing CSS evaluation specs for the axes. These currently fail due to the ~ and + axes not being evaluated properly.	2014-11-12 23:38:17 +01:00
Yorick Peterse	817a5e075b	Wrap predicate AST nodes _around_ other nodes. This means that "foo[1]" uses this AST: (predicate (test nil "foo") (int 1)) Instead of this AST: (test nil "foo" (int 1)) This makes it easier for the XPath evaluator to process predicates correctly.	2014-11-12 22:59:38 +01:00
Yorick Peterse	24350fa457	Added various predicate specs for XPath axes.	2014-11-12 09:40:22 +01:00
Yorick Peterse	c15604a86f	CSS evaluator specs for IDs.	2014-11-11 00:21:28 +01:00
Yorick Peterse	b9e1b51270	CSS evaluator specs for classes.	2014-11-11 00:18:44 +01:00
Yorick Peterse	43200238c5	CSS evaluator specs for predicates and operators.	2014-11-10 23:21:00 +01:00

1 2 3 4 5 ...

425 Commits